> This page is for version Latest · v1.4.0 (26.09) (default).
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.models.audio.sed.base

Stage-adapter contract for audio sound-event detection.

`SEDInferenceStage` owns Curator-side glue: reading `AudioTask.data`,
loading or normalizing audio, resampling, resume behavior, and writing output
fields or NPZ sidecars. `SEDAdapter` owns model construction, checkpoint
loading, model-specific batch padding, inference, and temporal metadata.

Keeping this boundary explicit lets a YAML pipeline replace PANNs with another
SED runtime by changing `adapter_target` while preserving the task schema.

## Module Contents

### Classes

| Name                                                           | Description                                                   |
| -------------------------------------------------------------- | ------------------------------------------------------------- |
| [`SEDAdapter`](#nemo_curator-models-audio-sed-base-SEDAdapter) | Structural protocol implemented by every sound-event adapter. |
| [`SEDResult`](#nemo_curator-models-audio-sed-base-SEDResult)   | Canonical result for one waveform returned by an SED adapter. |

### API

```python
class nemo_curator.models.audio.sed.base.SEDAdapter()
```

Protocol

Structural protocol implemented by every sound-event adapter.

Constructor contract: the stage creates an adapter as
`cls(checkpoint_path=..., sample_rate=..., **adapter_kwargs)`. A `None`
checkpoint path asks the adapter to resolve its registered default.

`infer_batch` receives stage-normalized items in input order. Each item
contains one contiguous mono float32 `waveform`. The adapter must return
exactly one `SEDResult` per item, in the same order.

**`checkpoint_path`** `str | None`

---

**`sample_rate`** `int`

---

```python
nemo_curator.models.audio.sed.base.SEDAdapter.download_weights_on_node() -> None
```

Cache model weights without allocating worker-local model state.

```python
nemo_curator.models.audio.sed.base.SEDAdapter.infer_batch(
    items: list[dict[str, typing.Any]]
) -> list[nemo_curator.models.audio.sed.base.SEDResult]
```

Return one canonical result per prepared waveform, in order.

```python
nemo_curator.models.audio.sed.base.SEDAdapter.load_model(
    num_gpus: int
) -> None
```

Load worker-local model state for the requested physical GPU count.

```python
nemo_curator.models.audio.sed.base.SEDAdapter.unload_model() -> None
```

Release worker-local model and accelerator state.

```python
class nemo_curator.models.audio.sed.base.SEDResult(
    framewise_output: numpy.ndarray,
    fps: float,
    valid_frames: int,
    original_num_samples: int
)
```

Dataclass

Canonical result for one waveform returned by an SED adapter.

**`fps`** `float`

Number of output frames per second.

---

**`framewise_output`** `ndarray`

Two-dimensional `(frames, classes)` probability
matrix. It may include a padded tail shared with other batch rows.

---

**`original_num_samples`** `int`

Real waveform length after stage resampling and
before model-specific padding.

---

**`valid_frames`** `int`

Number of leading rows that correspond to real audio.

---