nemo_curator.models.audio.sed.panns

View as Markdown

PANNs CNN14 implementation of the shared sound-event adapter.

The Curator stage supplies mono float32 waveforms at sample_rate. This adapter selects a checkpoint-compatible CNN14 variant, pads one ragged batch, runs one model call, and packages each row as SEDResult.

Module Contents

Classes

NameDescription
PANNsSEDAdapterRun an AudioSet-pretrained PANNs CNN14 checkpoint.

Data

_DEFAULT_CHECKPOINT_CONFIG

_DEFAULT_CHECKPOINT_FILENAME

_DEFAULT_CHECKPOINT_URL

_DEFAULT_MODEL_TYPE

API

class nemo_curator.models.audio.sed.panns.PANNsSEDAdapter(
checkpoint_path: str | None = None,
sample_rate: int = _DEFAULT_CHECKPOINT_CONFIG[...,
model_type: str = _DEFAULT_MODEL_TYPE,
window_size: int = _DEFAULT_CHECKPOINT_CONFIG[...,
hop_size: int = _DEFAULT_CHECKPOINT_CONFIG[...,
mel_bins: int = _DEFAULT_CHECKPOINT_CONFIG[...,
fmin: int = _DEFAULT_CHECKPOINT_CONFIG[...,
fmax: int = _DEFAULT_CHECKPOINT_CONFIG[...,
classes_num: int = _DEFAULT_CHECKPOINT_CONFIG[...,
pad_short_segments: bool = True
)
Dataclass

Run an AudioSet-pretrained PANNs CNN14 checkpoint.

model_type must be one of the checkpoint names exposed by nemo_curator.models.audio.sed.SUPPORTED_MODEL_TYPES. The frontend arguments must match the checkpoint. checkpoint_path is the single checkpoint location: an existing file is loaded directly, while a missing file path is populated from the upstream PANNs Zenodo release. With no path, that same existing-or-missing behavior applies to the standard PyTorch Hub checkpoint file. The default configuration produces 527-class output at 100 frames per second.

_device
Any = field(default=None, init=False, repr=False)
_model
Any = field(default=None, init=False, repr=False)
checkpoint_path
str | None = None
classes_num
int = _DEFAULT_CHECKPOINT_CONFIG['classes_num']
fmax
int = _DEFAULT_CHECKPOINT_CONFIG['fmax']
fmin
int = _DEFAULT_CHECKPOINT_CONFIG['fmin']
hop_size
int = _DEFAULT_CHECKPOINT_CONFIG['hop_size']
mel_bins
int = _DEFAULT_CHECKPOINT_CONFIG['mel_bins']
model_type
str = _DEFAULT_MODEL_TYPE
pad_short_segments
bool = True
sample_rate
int = _DEFAULT_CHECKPOINT_CONFIG['sample_rate']
window_size
int = _DEFAULT_CHECKPOINT_CONFIG['window_size']
nemo_curator.models.audio.sed.panns.PANNsSEDAdapter.__post_init__() -> None
nemo_curator.models.audio.sed.panns.PANNsSEDAdapter._download_default_checkpoint(
checkpoint_path: pathlib.Path
) -> dict[str, typing.Any]

Download or reuse the registered default at the requested file location.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter._load_checkpoint() -> dict[str, typing.Any]

Load an existing checkpoint or download the registered default.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter._pad_to_rectangle(
waveforms: list[numpy.ndarray]
) -> numpy.ndarray

Zero-pad a ragged batch to the CNN14 minimum and longest row.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter._resolve_checkpoint_path() -> pathlib.Path

Return the configured file or the standard PyTorch Hub checkpoint file.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter._validate_default_checkpoint_config() -> None

Reject incompatible settings before resolving the registered default.

For example, PANNsSEDAdapter(sample_rate=16_000).download_weights_on_node() reaches this validation when its resolved checkpoint file is missing and raises because the registered checkpoint uses its published 32 kHz frontend. Any existing resolved file, including one in the standard PyTorch Hub cache, bypasses this download-time validation.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter.download_weights_on_node() -> None

Ensure the checkpoint exists without constructing or loading the model.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter.infer_batch(
items: list[dict[str, typing.Any]]

Run one checkpoint-compatible CNN14 call for the prepared batch.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter.load_model(
num_gpus: int
) -> None

Load one single-device model on CPU or CUDA.

Like the direct audio GPU stages, zero selects CPU and any positive worker GPU allocation selects CUDA. PANNs remains single-device even when a worker reserves more than one GPU.

nemo_curator.models.audio.sed.panns.PANNsSEDAdapter.unload_model() -> None

Release the model and any reclaimable CUDA cache state.

nemo_curator.models.audio.sed.panns._DEFAULT_CHECKPOINT_CONFIG: dict[str, int] = {'sample_rate': 32000, 'window_size': 1024, 'hop_size': 320, 'mel_bins': 64, 'fm...
nemo_curator.models.audio.sed.panns._DEFAULT_CHECKPOINT_FILENAME = 'Cnn14_DecisionLevelMax_mAP=0.385.pth'
nemo_curator.models.audio.sed.panns._DEFAULT_CHECKPOINT_URL = 'https://zenodo.org/record/3987831/files/Cnn14_DecisionLevelMax_mAP%3D0.385.pth?...
nemo_curator.models.audio.sed.panns._DEFAULT_MODEL_TYPE = 'Cnn14_DecisionLevelMax'