NVIDIA Speech NIM Microservices Release Notes#
This page lists changes, fixes, and known issues for the current NVIDIA Speech NIM microservices release and links to previous release notes.
All Speech NIM microservice updates are released together as a collection and follow calendar versioning YY.MM.n, where YY is the year, MM is the month, and n is the patch number within that cycle.
Release 26.10.0#
NIM Container Versions#
NIM |
Container Tag |
|---|---|
Nemotron ASR Streaming |
Highlights#
Adds an eight-speaker Sortformer diarization model for Nemotron ASR Streaming and support for limiting the number of detected speaker identities independently for each stream.
Adds processed-audio duration usage to ASR HTTP JSON responses and realtime WebSocket transcription completion events.
Adds a
batch_size=512maximum-throughput profile for the English and multilingual Nemotron ASR Streaming model types.Nemotron ASR Streaming now supports approximately five times as many concurrent streams as version 1.3.1.
ASR NIM#
Key Features#
ASR HTTP JSON responses from
/v1/audio/transcriptionsnow containusage.type=durationandusage.seconds. The duration covers the complete audio processed for the request. Plain-text responses are unchanged.Realtime
conversation.item.input_audio_transcription.completedevents now contain cumulative processed-audio duration in theusageobject. The terminal event contains the total duration for the stream.Nemotron ASR Streaming now uses an eight-speaker Sortformer diarization model. Set
max_speaker_countfrom 1 through 8 to limit the number of detected speaker identities for an individual stream. The value is an upper bound, not an exact requested speaker count. For gRPC, omit the value or set it to0to use the eight-speaker default. Parakeet CTC and Parakeet RNNT profiles continue to use the four-speaker Sortformer model.Nemotron ASR Streaming now provides
batch_size=32,batch_size=128, andbatch_size=512profiles for bothtype=en-USandtype=multi. The newbatch_size=512profile uses an acoustic model batch size of 512 and a pipeline batch size of 4096, is optimized for maximum throughput, and requires a GPU with at least 48 GB of memory.Adds configurable beam search for cache-aware streaming RNNT models.
Changes word-boosting scores to the range
-200through200for CTC models and Nemotron ASR Streaming. Other RNNT and TDT models continue to support scores from0.5through2.0.
Known Issues#
Speaker diarization accuracy can decrease for recordings with more than four speakers. If you know the approximate number of speakers in the recording, set the client-side
max_speaker_countparameter to that number. For the sample client, use--diarization-max-speakers. Restricting the maximum number of speakers can reduce speaker confusion and improve speaker attribution. This value is an upper bound; the model can still identify fewer speakers.
Support Matrix and Compatibility Updates#
Updates Nemotron ASR Streaming profile selection and capacity information for the new
batch_size=512profile.Adds eight-speaker diarization support to Nemotron ASR Streaming. Parakeet CTC and Parakeet RNNT diarization remains limited to four speakers.
Previous Release Notes#
Archived Documentation#
With the introduction of the NVIDIA Speech NIM microservices documentation beginning with release 26.02.0, the previous NVIDIA Riva NIM documentation has been officially deprecated. To access the deprecated documentation, refer to the following links: