Beginner Audio Processing Tutorial
Run the maintained FLEURS tutorial to download multilingual speech, transcribe it with a NeMo ASR model, calculate Word Error Rate (WER) and duration, filter samples, and write a JSONL manifest.
What You Will Do
The checked-in tutorial builds this pipeline:
CreateInitialManifestFleursStagedownloads and prepares a FLEURS split.ASRStageandNeMoASRAdaptergeneratepred_text.GetPairwiseWerStagewriteswer_pct.GetAudioDurationStagewritesduration.PreserveByValueStagefilters by the configured WER threshold.AudioToDocumentStageandJsonlWriterwrite the filtered JSONL output.
Maintained Example Files
The working tutorial is maintained in one place:
main.py loads pipeline.yaml through Hydra and manages the Ray client and selected executor. Use this entrypoint instead of copying the pipeline into a second script.
Prerequisites
- Python 3.11, 3.12, or 3.13
- NeMo Curator installed with
audio_cuda12for GPU inference oraudio_cpufor CPU inference - Internet access for the FLEURS split and ASR checkpoint
- Sufficient disk space for the selected language and split
- An NVIDIA GPU with approximately 4 GB or more VRAM is recommended; CPU fallback is supported but substantially slower
From the repository root, install one audio extra and activate the environment:
Run the Default Pipeline
Run the maintained entrypoint from the repository root:
The defaults process the Armenian (hy_am) development split with nvidia/stt_hy_fastconformer_hybrid_large_pc, keep samples with wer_pct <= 5.5, use the Xenna backend, and write results under ${raw_data_dir}/result/${lang}/.
CPU fallback
Set the ASR stage’s GPU requirement to zero:
CPU inference is intended for functional testing and is typically 10–50 times slower than GPU inference.
Understand the ASR Capacity Contract
The supplied pipeline.yaml keeps the required ASR capacity settings together:
max_inference_duration_s must be less than or equal to max_audio_sec_per_actor. If you reduce the actor budget to address memory pressure, reduce the inference-duration ceiling as needed to preserve that invariant.
Customize the Run
Use Hydra overrides without editing pipeline.yaml:
Common overrides:
Refer to tutorials/audio/fleurs/pipeline.yaml for the complete configuration contract and tutorials/audio/fleurs/README.md for model, backend, and performance guidance.
Inspect the Results
For the default language, output is written under:
Each output row contains the top-level fields produced by the maintained pipeline:
Inspect the output from the repository root: