nemo_curator.models.audio.sed.base
nemo_curator.models.audio.sed.base
Stage-adapter contract for audio sound-event detection.
SEDInferenceStage owns Curator-side glue: reading AudioTask.data,
loading or normalizing audio, resampling, resume behavior, and writing output
fields or NPZ sidecars. SEDAdapter owns model construction, checkpoint
loading, model-specific batch padding, inference, and temporal metadata.
Keeping this boundary explicit lets a YAML pipeline replace PANNs with another
SED runtime by changing adapter_target while preserving the task schema.
Module Contents
Classes
API
Structural protocol implemented by every sound-event adapter.
Constructor contract: the stage creates an adapter as
cls(checkpoint_path=..., sample_rate=..., **adapter_kwargs). A None
checkpoint path asks the adapter to resolve its registered default.
infer_batch receives stage-normalized items in input order. Each item
contains one contiguous mono float32 waveform. The adapter must return
exactly one SEDResult per item, in the same order.
Cache model weights without allocating worker-local model state.
Return one canonical result per prepared waveform, in order.
Load worker-local model state for the requested physical GPU count.
Release worker-local model and accelerator state.
Canonical result for one waveform returned by an SED adapter.
Number of output frames per second.
Two-dimensional (frames, classes) probability
matrix. It may include a padded tail shared with other batch rows.
Real waveform length after stage resampling and before model-specific padding.
Number of leading rows that correspond to real audio.