> This page is for version Latest · v1.4.0 (26.09) (default).
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.models.audio.indic_conformer_hybrid

AI4Bharat IndicConformer *hybrid* (CTC+RNNT) per-language `.nemo` ASR.

This adapter loads the per-language
`ai4bharat/indicconformer_stt_&lt;lang&gt;_hybrid_ctc_rnnt_large` `.nemo`
checkpoints and runs waveform-to-text inference behind the shared `ASRStage`.

These checkpoints were trained with AI4Bharat's NeMo fork
([https://github.com/AI4Bharat/NeMo](https://github.com/AI4Bharat/NeMo), `nemo-v2` branch), which adds a *multi-softmax*
head to the standard NeMo ASR models: one shared Conformer encoder + shared RNNT
prediction network, and a **per-language output head** selected at inference time by
`language_id`.

The stock `nemo-toolkit` (2.7.x) installed in this container does NOT know those
config keys, so `ASRModel.restore_from` fails out of the box:

* `RNNTDecoder(multisoftmax=...)`      -> unexpected kwarg
* `RNNTJoint(multilingual=..., language_keys=...)` -> unexpected kwargs +
  a per-language `ModuleDict` final layer instead of a single `Linear`
* `ConvASRDecoder(multisoftmax=...)`   -> unexpected kwarg

Rather than installing the fork (which is pinned to NeMo 1.23 and would break the
rest of the pipeline), `_apply_multisoftmax_patches` **monkeypatches just those
three module classes** on top of the installed NeMo so the checkpoint loads, and the
model then runs a **compact greedy CTC / RNNT decode** that mirrors the fork's decode
semantics (per-language blank index `V/num_langs`, per-language joint head, local-id
feedback to the prediction network). Decoding maps the per-language local token ids
back to text through the model's own `AggregateTokenizer` (which already ships the
per-language tokenizers and offset tables in 2.7.x).

The patches are idempotent and additive: when `multisoftmax` / `multilingual` are
absent (a normal NeMo model), every patched path falls back to the original behaviour,
so importing this module does not change ordinary NeMo usage.

## Module Contents

### Classes

| Name                                                                                                   | Description                                                            |
| ------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------- |
| [`IndicConformerHybridASR`](#nemo_curator-models-audio-indic_conformer_hybrid-IndicConformerHybridASR) | AI4Bharat IndicConformer hybrid adapter for #1967's generic ASR stage. |
| [`_LanguageRNNTDecoder`](#nemo_curator-models-audio-indic_conformer_hybrid-_LanguageRNNTDecoder)       | Route a per-language blank to the aggregate predictor's SOS token.     |
| [`_LanguageRNNTJoint`](#nemo_curator-models-audio-indic_conformer_hybrid-_LanguageRNNTJoint)           | Bind the multilingual joint network to one language head.              |

### Functions

| Name                                                                                                           | Description                                                                    |
| -------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| [`_apply_multisoftmax_patches`](#nemo_curator-models-audio-indic_conformer_hybrid-_apply_multisoftmax_patches) | Idempotently patch ConvASRDecoder / RNNTJoint / RNNTDecoder for multi-softmax. |

### Data

[`INDIC_CONFORMER_HYBRID_LANGS`](#nemo_curator-models-audio-indic_conformer_hybrid-INDIC_CONFORMER_HYBRID_LANGS)

[`_JOINT_CTX`](#nemo_curator-models-audio-indic_conformer_hybrid-_JOINT_CTX)

[`_MAX_CHUNK_DURATION_SEC`](#nemo_curator-models-audio-indic_conformer_hybrid-_MAX_CHUNK_DURATION_SEC)

[`_PATCHED`](#nemo_curator-models-audio-indic_conformer_hybrid-_PATCHED)

[`_TARGET_SR`](#nemo_curator-models-audio-indic_conformer_hybrid-_TARGET_SR)

### API

```python
class nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR(
    model_id: str,
    revision: str | None = None,
    decode_mode: typing.Literal['ctc', 'rnnt'] = 'rnnt',
    max_symbols_per_step: int = 10,
    tensorrt_engine_dir: str | None = None,
    rnnt_precision: typing.Literal['fp32', 'fp16', 'bf16'] = 'fp32',
    empty_audio_marks_skip: bool = True
)
```

AI4Bharat IndicConformer hybrid adapter for #1967's generic ASR stage.

**`_chunk_duration_sec`** `float | None = _MAX_CHUNK_DURATION_SEC`

---

**`_num_langs`** `int = 0`

---

**`_per_lang_classes`** `int = 0`

---

**`_rnnt_decoders`** `dict[str, Any] = {}`

---

**`_trt_metadata`** `dict[str, Any] | None = None`

---

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._configure_rnnt_precision() -> None
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._decode_ctc_batch(
    encoded: typing.Any,
    encoded_len: typing.Any,
    lang_codes: list[str]
) -> list[str]
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._decode_ctc_row(
    log_probs: typing.Any,
    encoded_len: int,
    lang: str
) -> str
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._decode_encoded_batch(
    encoded: typing.Any,
    encoded_len: typing.Any,
    languages: list[str],
    mode: str
) -> list[str]
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._decode_rnnt_batch(
    encoded: typing.Any,
    encoded_len: typing.Any,
    lang_codes: list[str]
) -> list[str]
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._enable_tensorrt_encoder(
    engine_path: pathlib.Path
) -> None
```

Replace only the bundled NeMo model's encoder with TensorRT.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._existing_local_checkpoint(
    model_id: str
) -> str | None
```

staticmethod

Return an existing checkpoint file and reject local non-files.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._finalize_loaded_model() -> None
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._generate_chunks(
    waveforms: list[numpy.ndarray],
    sample_rates: list[int],
    lang_codes: list[str],
    mode: str
) -> tuple[list[str], list[str]]
```

Batch already bounded chunks by duration and restore chunk order.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._ids_to_text(
    local_ids: list[int],
    lang: str
) -> str
```

Map per-language local token ids -> aggregate ids -> text.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._offline() -> bool
```

staticmethod

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._resolve_nemo_path(
    model_id: str
) -> str
```

classmethod

Resolve a local checkpoint or an already-cached Hugging Face repo ID.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._resolve_tensorrt_bundle() -> tuple[dict[str, typing.Any], pathlib.Path, pathlib.Path]
```

Validate and resolve the three local TensorRT bundle artifacts.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._rnnt_decoder(
    lang: str
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR._rnnt_dtype() -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR.download_weights_on_node() -> None
```

Resolve the configured checkpoint into the node-local cache without loading it.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR.generate(
    waveforms: list[numpy.ndarray],
    sample_rates: list[int],
    lang_codes: list[str],
    decode_mode: str | None = None
) -> tuple[list[str], list[str]]
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR.load_model(
    num_gpus: int
) -> None
```

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR.transcribe_batch(
    items: list[dict[str, typing.Any]]
) -> list[nemo_curator.models.asr.base.ASRResult]
```

Transcribe supported rows and preserve the shared one-result-per-item contract.

```python
nemo_curator.models.audio.indic_conformer_hybrid.IndicConformerHybridASR.unload_model() -> None
```

```python
class nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTDecoder(
    decoder: typing.Any,
    blank_index: int
)
```

Route a per-language blank to the aggregate predictor's SOS token.

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTDecoder.__getattr__(
    name: str
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTDecoder.predict(
    y: typing.Any = None,
    state: typing.Any = None,
    kwargs: typing.Any = {}
) -> typing.Any
```

```python
class nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTJoint(
    joint: typing.Any,
    language: str,
    num_classes_with_blank: int
)
```

Bind the multilingual joint network to one language head.

**`num_classes_with_blank`** `int`

---

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTJoint.__getattr__(
    name: str
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTJoint.joint_after_projection(
    f: typing.Any,
    g: typing.Any
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTJoint.project_encoder(
    encoder_output: typing.Any
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._LanguageRNNTJoint.project_prednet(
    prednet_output: typing.Any
) -> typing.Any
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._apply_multisoftmax_patches() -> None
```

Idempotently patch ConvASRDecoder / RNNTJoint / RNNTDecoder for multi-softmax.

```python
nemo_curator.models.audio.indic_conformer_hybrid.INDIC_CONFORMER_HYBRID_LANGS: frozenset[str] = frozenset({'as', 'bn', 'brx', 'doi', 'gu', 'hi', 'kn', 'kok', 'ks', 'mai', 'ml',...
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._JOINT_CTX: dict[str, Any] = {}
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._MAX_CHUNK_DURATION_SEC = 40.0
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._PATCHED = False
```

```python
nemo_curator.models.audio.indic_conformer_hybrid._TARGET_SR = 16000
```