nemo_curator.models.audio.sed.panns
nemo_curator.models.audio.sed.panns
PANNs CNN14 implementation of the shared sound-event adapter.
The Curator stage supplies mono float32 waveforms at sample_rate. This
adapter selects a checkpoint-compatible CNN14 variant, pads one ragged batch,
runs one model call, and packages each row as SEDResult.
Module Contents
Classes
Data
API
Run an AudioSet-pretrained PANNs CNN14 checkpoint.
model_type must be one of the checkpoint names exposed by
nemo_curator.models.audio.sed.SUPPORTED_MODEL_TYPES. The frontend arguments
must match the checkpoint. checkpoint_path is the single checkpoint
location: an existing file is loaded directly, while a missing file path
is populated from the upstream PANNs Zenodo release. With no path, that
same existing-or-missing behavior applies to the standard PyTorch Hub
checkpoint file. The default configuration produces 527-class output at
100 frames per second.
Download or reuse the registered default at the requested file location.
Load an existing checkpoint or download the registered default.
Zero-pad a ragged batch to the CNN14 minimum and longest row.
Return the configured file or the standard PyTorch Hub checkpoint file.
Reject incompatible settings before resolving the registered default.
For example, PANNsSEDAdapter(sample_rate=16_000).download_weights_on_node()
reaches this validation when its resolved checkpoint file is missing and
raises because the registered checkpoint uses its published 32 kHz
frontend. Any existing resolved file, including one in the standard
PyTorch Hub cache, bypasses this download-time validation.
Ensure the checkpoint exists without constructing or loading the model.
Run one checkpoint-compatible CNN14 call for the prepared batch.
Load one single-device model on CPU or CUDA.
Like the direct audio GPU stages, zero selects CPU and any positive worker GPU allocation selects CUDA. PANNs remains single-device even when a worker reserves more than one GPU.
Release the model and any reclaimable CUDA cache state.