ASR Inference
Use ASRStage to transcribe audio with a supported model adapter.
Choose an Adapter
The maintained tutorials show complete manifest-to-transcript pipelines for these adapters:
Each tutorial includes its optional dependencies, executor settings, model defaults, and a runnable command. The examples use GPU inference unless the tutorial documents a CPU option.
Refer to NeMo ASR models for the NeMo model catalog and model-specific guidance.
Configure the Stage
Set adapter_target to the adapter import path and model_id to the model
checkpoint. The following example shows the shared stage interface:
Use the adapter-specific tutorial when you need adapter_kwargs. Those
settings stay with the adapter that implements them. Set resources.gpus to
the number of GPUs each worker receives.
Provide Input Audio
Each AudioTask must provide an audio path under the key configured by
audio_filepath_key. The default key is resampled_audio_filepath, so place a
ResampleAudioStage before ASRStage, or point the stage at an existing audio
path. For in-memory audio, set waveform_key and provide the sample rate under
sample_rate_key (default sampling_rate) instead. The stage normalizes either
input to mono at target_sample_rate (16 kHz by default) before inference.
For manifest pipelines, each JSONL row can contain an audio path and optional language metadata:
Set source_lang_key to use another language field. Set default_language
when rows do not contain language metadata. If you set
supported_language_codes, rows outside that list remain in the output with a
language skip reason instead of being sent to the adapter.
Read Output
The stage preserves input metadata and writes the transcript to pred_text_key,
which defaults to pred_text. With waveform input, the waveform is removed
after inference by default; set keep_waveform=True when a downstream stage
needs it. Set extras_key to store adapter-specific serializable metadata in
one nested field. The stage can also write these control fields:
Skipped results remain in the output with _skipme. By default, an unreadable
audio file receives an audio_load_error skip reason; set
fail_on_audio_error to True when that audio error must fail the pipeline.
Model loading and adapter inference errors propagate to the pipeline. Inspect
_skipme and additional_notes before consuming the output manifest.
Tune Batching and Resources
batch_size controls how many candidate rows the backend groups for a call to
process_batch. The stage splits recordings longer than
max_inference_duration_s into segments, then packs those segments into adapter
calls bounded by max_audio_sec_per_actor. Set
max_audio_sec_per_actor to control padded adapter-call memory; it must be at
least as large as max_inference_duration_s. Use local_bucketing=True to
group segments with similar durations and reduce padding. The stage stitches
segment transcripts into one result per input row. Pre-segment files only when
you need application-specific or semantic boundaries. Use the tutorial’s
documented CPU configuration when the selected adapter supports CPU inference.