Curate AudioProcess DataASR Inference

NeMo ASR Models

View as Markdown

Use NeMo Framework’s automatic speech recognition models for transcription in your audio curation pipelines. This guide covers basic usage and configuration.

Model Selection

NeMo Framework provides pre-trained ASR models through the Hugging Face model hub. For the complete list of available models and their specifications, refer to the NeMo Framework ASR documentation.

Example Model Usage

# Example using a test-verified model
example_model = "nvidia/parakeet-tdt-0.6b-v2"
# For production use, select appropriate models from:
# https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/asr/all_chkpt.html

Basic Usage

Simple ASR Inference

from nemo_curator.stages.audio.inference.asr.stage import ASRStage
from nemo_curator.stages.resources import Resources
# Create ASR inference stage with a model from NeMo Framework
asr_stage = ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/parakeet-tdt-0.6b-v2",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="audio_filepath",
pred_text_key="pred_text",
batch_size=16,
)
# Configure for GPU processing
asr_stage = asr_stage.with_(
resources=Resources(gpus=1.0),
)

Custom Configuration

# Example with custom field names
custom_asr = ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/parakeet-tdt-0.6b-v2",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="custom_audio_path",
pred_text_key="transcription",
batch_size=32,
).with_(
resources=Resources(cpus=4.0, gpus=1.0),
)

Model Caching

The executor downloads and caches model weights once per node, then loads one adapter in each worker:

asr_stage = ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/parakeet-tdt-0.6b-v2",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="audio_filepath",
)
# Add the stage to a Pipeline and run it through an executor. Do not call
# setup() or process_batch() directly; the executor owns worker lifecycle.

Resource Configuration

Configure GPU and CPU resources based on your hardware:

from nemo_curator.stages.resources import Resources
# Single GPU configuration
asr_stage = ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/parakeet-tdt-0.6b-v2",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="audio_filepath",
batch_size=16,
).with_(
resources=Resources(
cpus=4.0,
gpus=1.0,
),
)

NeMoASRAdapter uses zero or one GPU per worker. Scale throughput by allowing the executor to run more one-GPU workers instead of assigning multiple GPUs to one stage instance.

Resource requirements vary by model. Test with your specific model to determine optimal settings.