> This page is for version 26.07 · v1.3.0.
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.models.asr.indic_canary

Adapter for an Indic Canary model exported as a TensorRT-LLM engine.

## Module Contents

### Classes

| Name                                                                                 | Description                                                               |
| ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- |
| [`IndicCanaryTRTLLMASR`](#nemo_curator-models-asr-indic_canary-IndicCanaryTRTLLMASR) | Run static-batch Indic Canary inference from a prebuilt engine directory. |

### Data

[`_DEFAULT_MAX_DURATION_SEC`](#nemo_curator-models-asr-indic_canary-_DEFAULT_MAX_DURATION_SEC)

[`_DEFAULT_MIN_DURATION_SEC`](#nemo_curator-models-asr-indic_canary-_DEFAULT_MIN_DURATION_SEC)

[`_MIN_DURATION_SAMPLES`](#nemo_curator-models-asr-indic_canary-_MIN_DURATION_SAMPLES)

[`_REQUIRED_ENGINE_FILES`](#nemo_curator-models-asr-indic_canary-_REQUIRED_ENGINE_FILES)

[`_TARGET_SAMPLE_RATE`](#nemo_curator-models-asr-indic_canary-_TARGET_SAMPLE_RATE)

### API

```python
class nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR(
    engine_dir: str,
    num_beams: int = 4,
    max_new_tokens: int = 374,
    pnc: bool = False,
    max_duration_sec: float = _DEFAULT_MAX_DURATION_SEC,
    min_duration_sec: float = _DEFAULT_MIN_DURATION_SEC,
    kv_cache_free_gpu_memory_fraction: float = 0.2,
    cross_kv_cache_fraction: float = 0.2
)
```

Run static-batch Indic Canary inference from a prebuilt engine directory.

**`cross_kv_cache_fraction`** `= float(cross_kv_cache_fraction)`

---

**`kv_cache_free_gpu_memory_fraction`** `= float(kv_cache_free_gpu_memory_fraction)`

---

**`max_new_tokens`** `= int(max_new_tokens)`

---

**`max_samples`** `= int(self.max_duration_sec * _TARGET_SAMPLE_RATE)`

---

**`min_duration_sec`** `= min(min_duration_sec, self.max_duration_sec)`

---

**`min_samples`** `= int(self.min_duration_sec * _TARGET_SAMPLE_RATE)`

---

**`num_beams`** `= int(num_beams)`

---

**`pnc`** `= bool(pnc)`

---

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR._normalize_language(
    language: str
) -> str | None
```

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR._prompt_config(
    language: str
) -> dict[str, object]
```

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR._transcribe_prepared(
    padded: list[typing.Any],
    durations: list[int],
    prompts: list[dict[str, object]]
) -> list[str]
```

Run engine-sized sub-batches without exposing a second batch control.

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR.download_weights_on_node() -> None
```

Validate local engine artifacts without allocating GPU state.

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR.load_model(
    num_gpus: int
) -> None
```

Load the TensorRT-LLM runtime on its one required GPU.

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR.transcribe_batch(
    items: list[dict[str, typing.Any]]
) -> list[nemo_curator.models.asr.base.ASRResult]
```

Transcribe supported rows and preserve their original positions.

```python
nemo_curator.models.asr.indic_canary.IndicCanaryTRTLLMASR.unload_model() -> None
```

Release the engine and its CUDA allocations.

```python
nemo_curator.models.asr.indic_canary._DEFAULT_MAX_DURATION_SEC = 40.0
```

```python
nemo_curator.models.asr.indic_canary._DEFAULT_MIN_DURATION_SEC = 0.5
```

```python
nemo_curator.models.asr.indic_canary._MIN_DURATION_SAMPLES = 400
```

```python
nemo_curator.models.asr.indic_canary._REQUIRED_ENGINE_FILES = ('encoder/encoder.plan', 'encoder/config.json', 'decoder/config.json', 'decoder/...
```

```python
nemo_curator.models.asr.indic_canary._TARGET_SAMPLE_RATE = 16000
```