nemo_curator.models.audio.sed.base

View as Markdown

Stage-adapter contract for audio sound-event detection.

SEDInferenceStage owns Curator-side glue: reading AudioTask.data, loading or normalizing audio, resampling, resume behavior, and writing output fields or NPZ sidecars. SEDAdapter owns model construction, checkpoint loading, model-specific batch padding, inference, and temporal metadata.

Keeping this boundary explicit lets a YAML pipeline replace PANNs with another SED runtime by changing adapter_target while preserving the task schema.

Module Contents

Classes

NameDescription
SEDAdapterStructural protocol implemented by every sound-event adapter.
SEDResultCanonical result for one waveform returned by an SED adapter.

API

class nemo_curator.models.audio.sed.base.SEDAdapter()
Protocol

Structural protocol implemented by every sound-event adapter.

Constructor contract: the stage creates an adapter as cls(checkpoint_path=..., sample_rate=..., **adapter_kwargs). A None checkpoint path asks the adapter to resolve its registered default.

infer_batch receives stage-normalized items in input order. Each item contains one contiguous mono float32 waveform. The adapter must return exactly one SEDResult per item, in the same order.

checkpoint_path
str | None
sample_rate
int
nemo_curator.models.audio.sed.base.SEDAdapter.download_weights_on_node() -> None

Cache model weights without allocating worker-local model state.

nemo_curator.models.audio.sed.base.SEDAdapter.infer_batch(
items: list[dict[str, typing.Any]]

Return one canonical result per prepared waveform, in order.

nemo_curator.models.audio.sed.base.SEDAdapter.load_model(
num_gpus: int
) -> None

Load worker-local model state for the requested physical GPU count.

nemo_curator.models.audio.sed.base.SEDAdapter.unload_model() -> None

Release worker-local model and accelerator state.

class nemo_curator.models.audio.sed.base.SEDResult(
framewise_output: numpy.ndarray,
fps: float,
valid_frames: int,
original_num_samples: int
)
Dataclass

Canonical result for one waveform returned by an SED adapter.

fps
float

Number of output frames per second.

framewise_output
ndarray

Two-dimensional (frames, classes) probability matrix. It may include a padded tail shared with other batch rows.

original_num_samples
int

Real waveform length after stage resampling and before model-specific padding.

valid_frames
int

Number of leading rows that correspond to real audio.