Release Notes for NVIDIA NeMo Retriever Reranking NIM#
This documentation contains the release notes for NVIDIA NeMo Retriever Reranking NIM.
Note
Some releases are labelled “Production Branch” or “(PB)”. Production Branches provide reliable, stable versions of the NIM. Non-production branch releases (sometimes called Feature Branch (FB) releases) contain the latest features, improvements, and optimizations.
Release 2.3#
This release adds support for NVIDIA DGX Spark to the nvidia/llama-nemotron-rerank-vl-1b-v2 NIM.
Highlights#
Added FP16 support for NVIDIA GB10 on DGX Spark.
Added
aarch64/Arm64 container support for DGX Spark deployments.The standard single-GPU Docker deployment workflow now supports DGX Spark. For details, refer to Get Started.
For details, refer to Support Matrix.
Known Issues#
cuDNN SDPA is unusable on NVIDIA GB10 (SM121).
All Known Issues#
The known issues for NeMo Retriever Reranking NIM are the following:
For GPUs with less VRAM, such as A10G and L4, set
NIM_PIPELINE_MAX_BATCH_SIZEto 26 or lower. For details, refer to Environment Variables for NVIDIA NeMo Retriever Reranking NIM.
Release Notes for Previous Versions#
2.2 – Release 2.2 was intentionally skipped.
2.1 – Release 2.1 was intentionally skipped.