> This page is for version Main · preview.
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

> Run PANNs sound event detection on audio manifests with NeMo Curator

# Sound Event Detection

Run sound event detection (SED) on an audio manifest with NeMo Curator. The maintained [SED pipeline configuration](https://github.com/NVIDIA-NeMo/Curator/blob/main/tutorials/audio/sed/pipeline.yaml) runs the PANNs CNN14 model. It uses the AudioSet class vocabulary.

## Input Manifest

Create a JSONL file with one audio file per line. Each row must include `audio_filepath` with an absolute or resolvable path.

```json
{"audio_filepath": "/data/audio/example.wav"}
```

## Run the Pipeline

From the NeMo Curator repository root, install the audio CUDA extra and run the maintained YAML configuration:

```bash
uv run --extra audio_cuda12 python nemo_curator/config/run.py \
  --config-path ../../tutorials/audio/sed \
  --config-name pipeline \
  manifest_path=/absolute/path/to/input.jsonl \
  output_dir=/absolute/path/to/sed_output
```

The pipeline resamples audio to 32 kHz, converts it to mono, and runs GPU inference. It loads the default `Cnn14_DecisionLevelMax_mAP=0.385.pth` checkpoint from the PyTorch cache when present. When the checkpoint is missing, NeMo Curator downloads it automatically. Set `checkpoint_path` to use a specific file. NeMo Curator loads an existing file or downloads the default checkpoint to that path.

Use the decision-level checkpoint for this framewise SED pipeline. The `Cnn14_mAP=0.431.pth` checkpoint supports clip-level audio tagging and is not trained for framewise SED.

## Outputs

For each successful task, the stage adds the following fields:

* `_sed_framewise`: A `(frames, 527)` AudioSet probability matrix at approximately 100 frames per second.
* `sed_valid_frames`: The number of frames that correspond to the original audio before batch padding.
* `sed_fps`: The frame rate of the probability matrix.
* `npz_filepath`: The path to a compressed NPZ sidecar when `save_npz` is enabled, as it is in the maintained YAML.

The probability matrix contains model scores for each frame and class. The pipeline does not threshold scores or produce labeled event segments.

Each NPZ sidecar contains `framewise`, `fps`, `valid_frames`, `audio_filepath`, and `original_num_samples`. The pipeline writes sidecars below `<output_dir>/framewise`.