nemo_curator.models.asr.nemo_asr

View as Markdown

NeMo Framework ASR behind the shared ASRAdapter contract.

Module Contents

Classes

NameDescription
NeMoASRAdapterRun a pretrained NeMo checkpoint using waveforms prepared by ASRStage.

Functions

NameDescription
_extract_nemo_transcription_textsExtract text from the output shapes used by supported NeMo ASR models.
_nemo_asr_module-

API

class nemo_curator.models.asr.nemo_asr.NeMoASRAdapter(
model_id: str = _DEFAULT_FASTCONFORMER_CTC_...,
num_workers: int = 0,
empty_audio_marks_skip: bool = True,
verbose: bool = False,
enable_local_attention: bool = False,
local_attention_context_size: tuple[int, int] = (128, 128),
use_cuda_graph_decoder: bool | None = None,
refresh_cache: bool = False,
strict: bool = True
)
Dataclass

Run a pretrained NeMo checkpoint using waveforms prepared by ASRStage.

Parameters:

model_id
strDefaults to _DEFAULT_FASTCONFORMER_CTC_MODEL

Pretrained NeMo ASR checkpoint name.

num_workers
intDefaults to 0

Data-loader workers used by NeMo’s transcription call.

verbose
boolDefaults to False

Forward NeMo transcription progress output.

enable_local_attention
boolDefaults to False

Convert a compatible FastConformer checkpoint from global to local attention after loading.

local_attention_context_size
tuple[int, int]Defaults to (128, 128)

Left and right local-attention context.

use_cuda_graph_decoder
bool | NoneDefaults to None

Override NeMo’s RNNT CUDA-graph decoder. Leave as None to preserve the checkpoint default. Set to False on GPU/driver combinations that do not support NeMo’s label-loop CUDA graph implementation.

refresh_cache
boolDefaults to False

Forward NeMo’s checkpoint cache refresh flag.

strict
boolDefaults to True

Forward NeMo’s strict checkpoint loading flag.

_ATTENTION_CONTEXT_DIRECTIONS
int = 2
_DEFAULT_FASTCONFORMER_CTC_MODEL
str = 'nvidia/stt_en_fastconformer_ctc_large'
_DEFAULT_SAMPLE_RATE
int = 16000
_model
Any = field(default=None, init=False, repr=False)
empty_audio_marks_skip
bool = True
enable_local_attention
bool = False
local_attention_context_size
tuple[int, int] = (128, 128)
model_id
str = _DEFAULT_FASTCONFORMER_CTC_MODEL
num_workers
int = 0
refresh_cache
bool = False
strict
bool = True
use_cuda_graph_decoder
bool | None = None
verbose
bool = False
nemo_curator.models.asr.nemo_asr.NeMoASRAdapter.__post_init__() -> None
nemo_curator.models.asr.nemo_asr.NeMoASRAdapter._configure_local_attention(
model: typing.Any
) -> None
nemo_curator.models.asr.nemo_asr.NeMoASRAdapter._configure_rnnt_cuda_graph_decoder(
model: typing.Any
) -> None

Override the CUDA-graph setting on a compatible NeMo RNNT decoder.

nemo_curator.models.asr.nemo_asr.NeMoASRAdapter._load_checkpoint(
device: typing.Any
) -> typing.Any
nemo_curator.models.asr.nemo_asr.NeMoASRAdapter._transcribe_waveforms(
waveforms: list[numpy.ndarray]
) -> list[str]

Run one NeMo call for the batch already bounded by ASRStage.

nemo_curator.models.asr.nemo_asr.NeMoASRAdapter.download_weights_on_node() -> None

Download a pretrained checkpoint without allocating a GPU model.

nemo_curator.models.asr.nemo_asr.NeMoASRAdapter.load_model(
num_gpus: int
) -> None

Load one worker-local model on the device requested by ASRStage.

nemo_curator.models.asr.nemo_asr.NeMoASRAdapter.transcribe_batch(
items: list[dict[str, typing.Any]]

Transcribe one adapter call while preserving input order.

nemo_curator.models.asr.nemo_asr.NeMoASRAdapter.unload_model() -> None

Release worker-local model and CUDA cache state.

nemo_curator.models.asr.nemo_asr._extract_nemo_transcription_texts(
outputs: object
) -> list[str]

Extract text from the output shapes used by supported NeMo ASR models.

nemo_curator.models.asr.nemo_asr._nemo_asr_module() -> typing.Any