Audio Quickstart
Quickstart for NeMo Curator Audio Curation
Use this quickstart to run NeMo Curator on a FLEURS audio dataset and inspect transcription-quality results. For complete setup details, refer to the Installation Guide.
Prerequisites
Prepare the following environment before running the audio pipeline:
- Python 3.11, 3.12, or 3.13 on x86_64 Linux. Ubuntu 20.04 or 22.04 is recommended.
packaging22.0 or later and Git.uvfor package management.- An
ffmpegexecutable onPATHfor stages that resample or convert audio. The Python extras do not install it; audio workflows are easiest to reproduce in a locally built container with FFmpeg. For Python package installs, refer to FFmpeg installation options. - An NVIDIA Volta-class or newer GPU with compute capability 7.0 or later and CUDA 12 is recommended for ASR inference; the example also supports CPU execution.
Audio curation requires x86_64 Linux. The audio_cpu and audio_cuda12 extras omit dependencies for arm64/aarch64, including NeMo ASR and diarization. Use an amd64 host or container for ASR, diarization, and tagging workflows.
Quickstart Steps
Install Curator, run the maintained FLEURS pipeline, and inspect its output.
-
Install NeMo Curator. A locally built container is recommended for audio workflows so FFmpeg is available on every worker.
Build an Image (Recommended)
Source Installation
PyPI Installation
Follow Build an Image, then add FFmpeg using Add FFmpeg for Audio or Video. This packages the audio system dependency in the image used by each worker.
-
Open a repository checkout. If you do not already have one, clone the source revision that matches your Curator installation.
-
Run the maintained FLEURS pipeline from the repository root.
The checked-in configuration downloads the Armenian (
hy_am) development split, runs FastConformer ASR, calculates word error rate and duration, and filters atwer_pct <= 5.5.The configuration sets
max_audio_sec_per_actor=240.0,max_inference_duration_s=120.0,local_bucketing=true, anduse_cuda_graph_decoder=false. The 120-second inference limit is within the 240-second actor budget.For CPU execution, install the
audio_cpuextra and set the GPU resource request to zero. Refer to the Beginner Audio Processing Tutorial for supported overrides and details about the checked-in configuration. -
Inspect the generated files and JSONL records. The output includes
audio_filepath,text,pred_text,wer_pct, anddurationfields.
Next Steps
Use these guides to adapt the pipeline to your audio datasets: