nemo_curator.models.audio.indic_conformer_hybrid
nemo_curator.models.audio.indic_conformer_hybrid
AI4Bharat IndicConformer hybrid (CTC+RNNT) per-language .nemo ASR.
This adapter loads the per-language
ai4bharat/indicconformer_stt_<lang>_hybrid_ctc_rnnt_large .nemo
checkpoints and runs waveform-to-text inference behind the shared ASRStage.
These checkpoints were trained with AI4Bharat’s NeMo fork
(https://github.com/AI4Bharat/NeMo, nemo-v2 branch), which adds a multi-softmax
head to the standard NeMo ASR models: one shared Conformer encoder + shared RNNT
prediction network, and a per-language output head selected at inference time by
language_id.
The stock nemo-toolkit (2.7.x) installed in this container does NOT know those
config keys, so ASRModel.restore_from fails out of the box:
RNNTDecoder(multisoftmax=...)-> unexpected kwargRNNTJoint(multilingual=..., language_keys=...)-> unexpected kwargs + a per-languageModuleDictfinal layer instead of a singleLinearConvASRDecoder(multisoftmax=...)-> unexpected kwarg
Rather than installing the fork (which is pinned to NeMo 1.23 and would break the
rest of the pipeline), _apply_multisoftmax_patches monkeypatches just those
three module classes on top of the installed NeMo so the checkpoint loads, and the
model then runs a compact greedy CTC / RNNT decode that mirrors the fork’s decode
semantics (per-language blank index V/num_langs, per-language joint head, local-id
feedback to the prediction network). Decoding maps the per-language local token ids
back to text through the model’s own AggregateTokenizer (which already ships the
per-language tokenizers and offset tables in 2.7.x).
The patches are idempotent and additive: when multisoftmax / multilingual are
absent (a normal NeMo model), every patched path falls back to the original behaviour,
so importing this module does not change ordinary NeMo usage.
Module Contents
Classes
Functions
Data
API
AI4Bharat IndicConformer hybrid adapter for #1967’s generic ASR stage.
Replace only the bundled NeMo model’s encoder with TensorRT.
Return an existing checkpoint file and reject local non-files.
Batch already bounded chunks by duration and restore chunk order.
Map per-language local token ids -> aggregate ids -> text.
Resolve a local checkpoint or an already-cached Hugging Face repo ID.
Validate and resolve the three local TensorRT bundle artifacts.
Resolve the configured checkpoint into the node-local cache without loading it.
Transcribe supported rows and preserve the shared one-result-per-item contract.
Route a per-language blank to the aggregate predictor’s SOS token.
Bind the multilingual joint network to one language head.
Idempotently patch ConvASRDecoder / RNNTJoint / RNNTDecoder for multi-softmax.