nemo_curator.models.audio.sed.cnn14
nemo_curator.models.audio.sed.cnn14
CNN14 decision-level model implementations for Sound Event Detection.
Vendored from the PANNs reference implementation at immutable commit
d2f4b8c18eab44737fcc0de1248ae21eb43f6aa4 (MIT, copyright 2018-2020
Qiuqiang Kong), whose license is reproduced above. The model blocks and three
decision-level CNN14 variants come from:
https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/d2f4b8c18eab44737fcc0de1248ae21eb43f6aa4/pytorch/models.py
The interpolation and frame-padding helpers come from:
https://github.com/qiuqiangkong/audioset_tagging_cnn/blob/d2f4b8c18eab44737fcc0de1248ae21eb43f6aa4/pytorch/pytorch_utils.py
See Kong et al., “PANNs: Large-Scale Pretrained Audio Neural Networks for
Audio Pattern Recognition” (2020).
Only the three decision-level variants are included because SED needs
framewise output; the base Cnn14 emits clip-level output only. The code
adds typing and factors shared helpers while retaining the reference forward
call contract; the architecture and tensor semantics are unchanged, so
published checkpoints load as-is.
Requires torchlibrosa, which arrives with the audio_cuda12 extra.
Module Contents
Classes
Functions
Data
API
Bases: Module
Bases: Module
CNN14 with decision-level attention for SED.
Bases: Module
CNN14 with decision-level average-pooling for SED.
Bases: Module
CNN14 with decision-level max-pooling for SED.
Bases: Module
Build the shared front-end layers for all CNN14 variants.
Run shared CNN14 encoding to get feature maps. Returns (features, frames_num).
Initialize a BatchNorm layer.
Initialize a Linear or Conv layer.
Interpolate in time to compensate CNN downsampling.
Returns: (batch, time_steps * ratio, classes_num)
Parameters:
(batch, time_steps, classes_num)
upsample factor
Pad framewise output to match input frame count.