NVIDIA Speech NIM Microservices Release Notes#

This page lists changes, fixes, and known issues for the current NVIDIA Speech NIM microservices release and links to previous release notes.

All Speech NIM microservice updates are released together as a collection and follow calendar versioning YY.MM.n, where YY is the year, MM is the month, and n is the patch number within that cycle.


Release 26.10.0#

NIM Container Versions#

NIM

Container Tag

Nemotron ASR Streaming

nemotron-asr-streaming:1.4.0

Highlights#

  • Adds an eight-speaker Sortformer diarization model for Nemotron ASR Streaming and support for limiting the number of detected speaker identities independently for each stream.

  • Adds processed-audio duration usage to ASR HTTP JSON responses and realtime WebSocket transcription completion events.

  • Adds a batch_size=512 maximum-throughput profile for the English and multilingual Nemotron ASR Streaming model types.

  • Nemotron ASR Streaming now supports approximately five times as many concurrent streams as version 1.3.1.

ASR NIM#

Key Features#

  • ASR HTTP JSON responses from /v1/audio/transcriptions now contain usage.type=duration and usage.seconds. The duration covers the complete audio processed for the request. Plain-text responses are unchanged.

  • Realtime conversation.item.input_audio_transcription.completed events now contain cumulative processed-audio duration in the usage object. The terminal event contains the total duration for the stream.

  • Nemotron ASR Streaming now uses an eight-speaker Sortformer diarization model. Set max_speaker_count from 1 through 8 to limit the number of detected speaker identities for an individual stream. The value is an upper bound, not an exact requested speaker count. For gRPC, omit the value or set it to 0 to use the eight-speaker default. Parakeet CTC and Parakeet RNNT profiles continue to use the four-speaker Sortformer model.

  • Nemotron ASR Streaming now provides batch_size=32, batch_size=128, and batch_size=512 profiles for both type=en-US and type=multi. The new batch_size=512 profile uses an acoustic model batch size of 512 and a pipeline batch size of 4096, is optimized for maximum throughput, and requires a GPU with at least 48 GB of memory.

  • Adds configurable beam search for cache-aware streaming RNNT models.

  • Changes word-boosting scores to the range -200 through 200 for CTC models and Nemotron ASR Streaming. Other RNNT and TDT models continue to support scores from 0.5 through 2.0.

Known Issues#

  • Speaker diarization accuracy can decrease for recordings with more than four speakers. If you know the approximate number of speakers in the recording, set the client-side max_speaker_count parameter to that number. For the sample client, use --diarization-max-speakers. Restricting the maximum number of speakers can reduce speaker confusion and improve speaker attribution. This value is an upper bound; the model can still identify fewer speakers.

Support Matrix and Compatibility Updates#

  • Updates Nemotron ASR Streaming profile selection and capacity information for the new batch_size=512 profile.

  • Adds eight-speaker diarization support to Nemotron ASR Streaming. Parakeet CTC and Parakeet RNNT diarization remains limited to four speakers.


Previous Release Notes#


Archived Documentation#

With the introduction of the NVIDIA Speech NIM microservices documentation beginning with release 26.02.0, the previous NVIDIA Riva NIM documentation has been officially deprecated. To access the deprecated documentation, refer to the following links: