About NeMo CuratorConceptsAudio Concepts

AudioTask Data Structure

View as Markdown

This guide covers the AudioTask data structure, which serves as the core container for audio data throughout NeMo Curator’s audio processing pipeline.

Overview

AudioTask is a specialized data structure that extends NeMo Curator’s base Task class to handle audio-specific processing requirements. Each AudioTask holds a single manifest entry, matching the convention used by VideoTask and FileGroupTask:

  • Single-Entry Model: One manifest entry per task (Task[dict]), enabling straightforward per-sample processing
  • Optional File Path Validation: validate() checks file existence only when filepath_key is configured and that key is present in the task data
  • Metadata Handling: Preserves audio characteristics and processing results throughout pipeline stages

Structure and Components

Basic Structure

from nemo_curator.tasks import AudioTask
# Create AudioTask with a single audio file
audio_task = AudioTask(
data={
"audio_filepath": "/path/to/audio.wav",
"text": "ground truth transcription",
"duration": 3.2,
"language": "en"
},
filepath_key="audio_filepath",
dataset_name="my_speech_dataset"
)

Key Attributes

AttributeTypeDescription
datadictAudio manifest entry (single dict, exposed as _AttrDict for attribute-style access)
filepath_keystr | NoneKey name for audio file paths in data (optional)
task_idstrFramework-managed deterministic lineage ID. Treat as read-only; it is assigned at stage boundaries.
dataset_namestrName of the source dataset
num_itemsintAlways returns 1 (read-only property)

Attribute-Style Access

AudioTask.data is an _AttrDict subclass, so you can access fields as attributes:

audio_task = AudioTask(data={"audio_filepath": "/path/to/audio.wav", "duration": 3.2})
# Both access styles work
audio_task.data["audio_filepath"] # dict-style
audio_task.data.audio_filepath # attribute-style

Data Validation

Explicit Validation

AudioTask provides an explicit validate() method. When filepath_key is set and that key is present in the task data, calling this method checks that the referenced local path exists. It does not reject a missing configured key. Constructing an AudioTask does not call validate() automatically, and tasks emitted by ManifestReader do not set filepath_key.

Metadata Management

Standard Metadata Fields

Common fields stored in AudioTask data:

audio_sample = {
# Core fields (user-provided)
"audio_filepath": "/path/to/audio.wav",
"text": "transcription text",
# Fields added by processing stages
"pred_text": "asr prediction", # Added by ASR inference stages
"wer_pct": 12.5, # Added by GetPairwiseWerStage
"duration": 3.2, # Added by GetAudioDurationStage
# Optional user-provided metadata
"language": "en_us",
"speaker_id": "speaker_001",
# Custom fields (examples)
"domain": "conversational",
"noise_level": "low"
}

Character error rate (CER) is available as a utility function and typically requires a custom stage to compute and store it.

Error Handling

Graceful Failure Modes

AudioTask handles various error conditions:

# Missing files
audio_task = AudioTask(data={
"audio_filepath": "/missing/file.wav", "text": "sample"
}, filepath_key="audio_filepath")
if not audio_task.validate():
print("Audio file is missing")
# Corrupted audio files
corrupted_sample = {
"audio_filepath": "/corrupted/audio.wav",
"text": "sample text"
}
# Duration calculation returns -1.0 for corrupted files
# Invalid metadata
invalid_sample = {
"audio_filepath": "/valid/audio.wav",
# Missing "text" field - needed for WER calculation but not enforced by AudioTask
}
# AudioTask does not enforce metadata field requirements. Add a validation stage if required.

Performance Characteristics

Memory Usage

AudioTask memory footprint is minimal since each task holds a single manifest entry. Memory scales with the number of metadata fields per entry and the total number of tasks processed in the pipeline.

Processing Patterns

Audio stages use several processing patterns:

PatternStagesMethod
Per-taskCPU stages (GetAudioDurationStage, GetPairwiseWerStage)process(task) → AudioTask — mutates task.data in-place
Batched audio transform/filterASRStage, PreserveByValueStageprocess_batch(tasks) → list[AudioTask]
Batched conversionAudioToDocumentStageprocess_batch(tasks) → list[DocumentBatch]

Integration with Processing Stages

Stage Input/Output

AudioTask serves as input and output for most audio transform and filter stages, which subclass ProcessingStage[AudioTask, AudioTask]. Source and conversion stages can use different task types; for example, AudioToDocumentStage is ProcessingStage[AudioTask, DocumentBatch].

# CPU stage: mutates task in-place and returns it
def process(self, task: AudioTask) -> AudioTask:
duration = get_duration(task.data["audio_filepath"])
task.data["duration"] = duration
return task

Chaining Stages

AudioTask flows through multiple processing stages, with each stage adding new metadata fields:

AudioTask (raw)• audio_filepath• text ASR Inference Stage AudioTask (with predictions)• audio_filepath• text• pred_text Quality Assessment Stage AudioTask (with metrics)• audio_filepath• text• pred_text• wer_pct• duration Filter Stage AudioTask (filtered)• audio_filepath• text• pred_text• wer_pct• duration Export Stage Output Files