Get Started

Audio Quickstart

View as Markdown

Quickstart for NeMo Curator Audio Curation

Use this quickstart to run NeMo Curator on a FLEURS audio dataset and inspect transcription-quality results. For complete setup details, refer to the Installation Guide.

Prerequisites

Prepare the following environment before running the audio pipeline:

  • Python 3.11, 3.12, or 3.13 on x86_64 Linux. Ubuntu 20.04 or 22.04 is recommended.
  • packaging 22.0 or later and Git.
  • uv for package management.
  • An ffmpeg executable on PATH for stages that resample or convert audio. The Python extras do not install it; audio workflows are easiest to reproduce in a locally built container with FFmpeg. For Python package installs, refer to FFmpeg installation options.
  • An NVIDIA Volta-class or newer GPU with compute capability 7.0 or later and CUDA 12 is recommended for ASR inference; the example also supports CPU execution.

Audio curation requires x86_64 Linux. The audio_cpu and audio_cuda12 extras omit dependencies for arm64/aarch64, including NeMo ASR and diarization. Use an amd64 host or container for ASR, diarization, and tagging workflows.

Quickstart Steps

Install Curator, run the maintained FLEURS pipeline, and inspect its output.

  1. Install NeMo Curator. A locally built container is recommended for audio workflows so FFmpeg is available on every worker.

  2. Open a repository checkout. If you do not already have one, clone the source revision that matches your Curator installation.

    if [ ! -d .git ]; then
    git clone https://github.com/NVIDIA-NeMo/Curator.git
    cd Curator
    fi
  3. Run the maintained FLEURS pipeline from the repository root.

    python tutorials/audio/fleurs/main.py \
    --config-path . \
    --config-name pipeline \
    raw_data_dir="${PWD}/example_audio/fleurs"

    The checked-in configuration downloads the Armenian (hy_am) development split, runs FastConformer ASR, calculates word error rate and duration, and filters at wer_pct <= 5.5.

    The configuration sets max_audio_sec_per_actor=240.0, max_inference_duration_s=120.0, local_bucketing=true, and use_cuda_graph_decoder=false. The 120-second inference limit is within the 240-second actor budget.

    For CPU execution, install the audio_cpu extra and set the GPU resource request to zero. Refer to the Beginner Audio Processing Tutorial for supported overrides and details about the checked-in configuration.

    python tutorials/audio/fleurs/main.py \
    --config-path . \
    --config-name pipeline \
    raw_data_dir="${PWD}/example_audio/fleurs-cpu" \
    stages.1.resources.gpus=0
  4. Inspect the generated files and JSONL records. The output includes audio_filepath, text, pred_text, wer_pct, and duration fields.

    example_audio/fleurs/
    ├── hy_am/ # Downloaded Armenian data and source manifest
    └── result/
    └── hy_am/ # Filtered JSONL output shards
    {
    "audio_filepath": "/absolute/path/to/audio.wav",
    "text": "ground truth transcription",
    "pred_text": "asr model prediction",
    "wer_pct": 2.5,
    "duration": 4.2
    }

Next Steps

Use these guides to adapt the pipeline to your audio datasets: