Curate AudioProcess DataASR Inference

ASR Inference

View as Markdown

Use ASRStage to transcribe audio with a supported model adapter.

Choose an Adapter

The maintained tutorials show complete manifest-to-transcript pipelines for these adapters:

AdapterUse WhenTutorial
NeMoASRAdapterYou need an NVIDIA NeMo FastConformer checkpoint.NeMo FastConformer ASR tutorial
QwenASRAdapterYou need Qwen3-ASR transcription with language routing.Qwen3-ASR adapter tutorial
QwenOmniASRAdapterYou need Qwen3-Omni in-process inference.Qwen3-Omni in-process ASR tutorial
FasterWhisperASRYou need Faster-Whisper Large-v3 transcription.Faster-Whisper adapter tutorial

Each tutorial includes its optional dependencies, executor settings, model defaults, and a runnable command. The examples use GPU inference unless the tutorial documents a CPU option.

Refer to NeMo ASR models for the NeMo model catalog and model-specific guidance.

Configure the Stage

Set adapter_target to the adapter import path and model_id to the model checkpoint. The following example shows the shared stage interface:

from nemo_curator.stages.audio.inference.asr.stage import ASRStage
from nemo_curator.stages.resources import Resources
asr = ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/stt_en_fastconformer_ctc_large",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="audio_filepath",
pred_text_key="pred_text",
batch_size=16,
resources=Resources(gpus=1.0),
)

Use the adapter-specific tutorial when you need adapter_kwargs. Those settings stay with the adapter that implements them. Set resources.gpus to the number of GPUs each worker receives.

Provide Input Audio

Each AudioTask must provide an audio path under the key configured by audio_filepath_key. The default key is resampled_audio_filepath, so place a ResampleAudioStage before ASRStage, or point the stage at an existing audio path. For in-memory audio, set waveform_key and provide the sample rate under sample_rate_key (default sampling_rate) instead. The stage normalizes either input to mono at target_sample_rate (16 kHz by default) before inference.

For manifest pipelines, each JSONL row can contain an audio path and optional language metadata:

{"audio_filepath": "/data/sample.wav", "source_lang": "en"}

Set source_lang_key to use another language field. Set default_language when rows do not contain language metadata. If you set supported_language_codes, rows outside that list remain in the output with a language skip reason instead of being sent to the adapter.

Read Output

The stage preserves input metadata and writes the transcript to pred_text_key, which defaults to pred_text. With waveform input, the waveform is removed after inference by default; set keep_waveform=True when a downstream stage needs it. Set extras_key to store adapter-specific serializable metadata in one nested field. The stage can also write these control fields:

FieldMeaning
pred_textTranscript text; skipped results can contain an empty string.
_skipmeMachine-readable reason for a skipped row.
additional_notesAdditional stage or adapter diagnostics.
Configured extras_keyAdapter-specific metadata when the adapter returns it.

Skipped results remain in the output with _skipme. By default, an unreadable audio file receives an audio_load_error skip reason; set fail_on_audio_error to True when that audio error must fail the pipeline. Model loading and adapter inference errors propagate to the pipeline. Inspect _skipme and additional_notes before consuming the output manifest.

Tune Batching and Resources

batch_size controls how many candidate rows the backend groups for a call to process_batch. The stage splits recordings longer than max_inference_duration_s into segments, then packs those segments into adapter calls bounded by max_audio_sec_per_actor. Set max_audio_sec_per_actor to control padded adapter-call memory; it must be at least as large as max_inference_duration_s. Use local_bucketing=True to group segments with similar durations and reduce padding. The stage stitches segment transcripts into one result per input row. Pre-segment files only when you need application-specific or semantic boundaries. Use the tutorial’s documented CPU configuration when the selected adapter supports CPU inference.