> This page is for version Latest · v1.4.0 (26.09) (default).
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.models.audio.sed.tensorrt

TensorRT execution for the CNN14 SED neural core.

## Module Contents

### Classes

| Name                                                                                         | Description                                                                |
| -------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
| [`SedCore`](#nemo_curator-models-audio-sed-tensorrt-SedCore)                                 | CNN14 neural core; the checkpoint's spectrogram frontend stays in PyTorch. |
| [`TensorRTPANNsSEDAdapter`](#nemo_curator-models-audio-sed-tensorrt-TensorRTPANNsSEDAdapter) | Run the PANNs CNN14 neural core with a target-specific TensorRT engine.    |
| [`TensorRTRunner`](#nemo_curator-models-audio-sed-tensorrt-TensorRTRunner)                   | Persistent TensorRT runner with shape-specific context memory.             |
| [`TensorRTSed`](#nemo_curator-models-audio-sed-tensorrt-TensorRTSed)                         | Reusable SED adapter preserving the checkpoint's PyTorch frontend.         |

### Functions

| Name                                                                                             | Description                                                                |
| ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------- |
| [`_sha256`](#nemo_curator-models-audio-sed-tensorrt-_sha256)                                     | -                                                                          |
| [`_trt_dtype_to_torch`](#nemo_curator-models-audio-sed-tensorrt-_trt_dtype_to_torch)             | -                                                                          |
| [`_validate_engine_metadata`](#nemo_curator-models-audio-sed-tensorrt-_validate_engine_metadata) | Reject an engine whose immutable build contract differs from this adapter. |
| [`extract_features`](#nemo_curator-models-audio-sed-tensorrt-extract_features)                   | Run the checkpoint's exact spectrogram and log-mel frontend.               |
| [`postprocess`](#nemo_curator-models-audio-sed-tensorrt-postprocess)                             | Restore PANNs framewise output geometry from CNN14 segment outputs.        |

### Data

[`_TENSORRT_MODEL_TYPE`](#nemo_curator-models-audio-sed-tensorrt-_TENSORRT_MODEL_TYPE)

### API

```python
class nemo_curator.models.audio.sed.tensorrt.SedCore(
    model: torch.nn.Module
)
```

**Bases:** `Module`

CNN14 neural core; the checkpoint's spectrogram frontend stays in PyTorch.

**`bn0`** `= model.bn0`

---

**`conv_block1`** `= model.conv_block1`

---

**`conv_block2`** `= model.conv_block2`

---

**`conv_block3`** `= model.conv_block3`

---

**`conv_block4`** `= model.conv_block4`

---

**`conv_block5`** `= model.conv_block5`

---

**`conv_block6`** `= model.conv_block6`

---

**`fc1`** `= model.fc1`

---

**`fc_audioset`** `= model.fc_audioset`

---

```python
nemo_curator.models.audio.sed.tensorrt.SedCore.forward(
    logmel: torch.Tensor
) -> torch.Tensor
```

```python
class nemo_curator.models.audio.sed.tensorrt.TensorRTPANNsSEDAdapter(
    checkpoint_path: str | None = None,
    sample_rate: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    model_type: str = _DEFAULT_MODEL_TYPE,
    window_size: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    hop_size: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    mel_bins: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    fmin: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    fmax: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    classes_num: int = _DEFAULT_CHECKPOINT_CONFIG[...,
    pad_short_segments: bool = True,
    tensorrt_engine_path: str | None = None,
    max_duration_sec: float | None = None
)
```

Dataclass

**Bases:** [PANNsSEDAdapter](/nemo/curator/nemo-curator/nemo_curator/models/audio/sed/panns#nemo_curator-models-audio-sed-panns-PANNsSEDAdapter)

Run the PANNs CNN14 neural core with a target-specific TensorRT engine.

Audio preprocessing, checkpoint resolution, batch padding, and the
canonical `SEDResult` contract match `PANNsSEDAdapter`. The checkpoint's
spectrogram and log-mel frontend remains in PyTorch; only the CNN14 neural
core runs in TensorRT. Engines are valid only for the GPU compute capability
and TensorRT version recorded in their adjacent JSON sidecar.

**`_max_input_samples`** `int | None = field(default=None, init=False, repr=False)`

---

**`max_duration_sec`** `float | None = None`

---

**`tensorrt_engine_path`** `str | None = None`

---

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTPANNsSEDAdapter.__post_init__() -> None
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTPANNsSEDAdapter.infer_batch(
    items: list[dict[str, object]]
) -> list[nemo_curator.models.audio.sed.base.SEDResult]
```

Run one TensorRT call and preserve the PANNs adapter result schema.

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTPANNsSEDAdapter.load_model(
    num_gpus: int
) -> None
```

Load the PyTorch frontend and TensorRT runtime on one CUDA device.

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTPANNsSEDAdapter.unload_model() -> None
```

Release the TensorRT context, PyTorch frontend, and CUDA cache.

```python
class nemo_curator.models.audio.sed.tensorrt.TensorRTRunner(
    engine_path: str | pathlib.Path,
    expected_metadata: collections.abc.Mapping[str, object]
)
```

Persistent TensorRT runner with shape-specific context memory.

**`_context`**

---

**`_device_memory`** `Tensor | None = None`

---

**`_engine`**

---

**`_input_names`** `list[str] = []`

---

**`_metadata`** `= self._validate_target(path, expected_metadata)`

---

**`_output_names`** `list[str] = []`

---

**`_runtime`** `= trt.Runtime(logger)`

---

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner.__call__(
    inputs: torch.Tensor = {}
) -> dict[str, torch.Tensor]
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner._bind_inputs(
    inputs: collections.abc.Mapping[str, torch.Tensor]
) -> torch.device
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner._prepare_device_memory(
    device: torch.device
) -> None
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner._validate_engine_io() -> None
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner._validate_target(
    path: pathlib.Path,
    expected_metadata: collections.abc.Mapping[str, object]
) -> dict[str, object]
```

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTRunner.close() -> None
```

```python
class nemo_curator.models.audio.sed.tensorrt.TensorRTSed(
    model: torch.nn.Module,
    engine_path: str | pathlib.Path,
    expected_metadata: collections.abc.Mapping[str, object]
)
```

Reusable SED adapter preserving the checkpoint's PyTorch frontend.

**`logmel`** `= model.logmel_extractor.to('cuda').eval()`

---

**`max_input_frames`** `int`

Maximum log-mel frame count accepted by the engine profile.

---

**`runner`**

---

**`spectrogram`** `= model.spectrogram_extractor.to('cuda').eval()`

---

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTSed.__call__(
    waveforms: torch.Tensor
) -> torch.Tensor
```

Return framewise probabilities for padded `[batch, samples]` input.

```python
nemo_curator.models.audio.sed.tensorrt.TensorRTSed.close() -> None
```

```python
nemo_curator.models.audio.sed.tensorrt._sha256(
    path: pathlib.Path
) -> str
```

```python
nemo_curator.models.audio.sed.tensorrt._trt_dtype_to_torch(
    dtype: object
) -> torch.dtype
```

```python
nemo_curator.models.audio.sed.tensorrt._validate_engine_metadata(
    engine_path: pathlib.Path,
    expected: collections.abc.Mapping[str, object],
    compute_capability: list[int],
    tensorrt_version: str
) -> dict[str, object]
```

Reject an engine whose immutable build contract differs from this adapter.

```python
nemo_curator.models.audio.sed.tensorrt.extract_features(
    model: torch.nn.Module,
    waveforms: torch.Tensor
) -> tuple[torch.Tensor, int]
```

Run the checkpoint's exact spectrogram and log-mel frontend.

```python
nemo_curator.models.audio.sed.tensorrt.postprocess(
    segmentwise: torch.Tensor,
    frames_num: int
) -> torch.Tensor
```

Restore PANNs framewise output geometry from CNN14 segment outputs.

```python
nemo_curator.models.audio.sed.tensorrt._TENSORRT_MODEL_TYPE = 'Cnn14_DecisionLevelMax'
```