AudioTask Data Structure
This guide covers the AudioTask data structure, which serves as the core container for audio data throughout NeMo Curator’s audio processing pipeline.
Overview
AudioTask is a specialized data structure that extends NeMo Curator’s base Task class to handle audio-specific processing requirements. Each AudioTask holds a single manifest entry, matching the convention used by VideoTask and FileGroupTask:
- Single-Entry Model: One manifest entry per task (
Task[dict]), enabling straightforward per-sample processing - Optional File Path Validation:
validate()checks file existence only whenfilepath_keyis configured and that key is present in the task data - Metadata Handling: Preserves audio characteristics and processing results throughout pipeline stages
Structure and Components
Basic Structure
Key Attributes
Attribute-Style Access
AudioTask.data is an _AttrDict subclass, so you can access fields as attributes:
Data Validation
Explicit Validation
AudioTask provides an explicit validate() method. When filepath_key is set and that key is present in the task data, calling this method checks that the referenced local path exists. It does not reject a missing configured key. Constructing an AudioTask does not call validate() automatically, and tasks emitted by ManifestReader do not set filepath_key.
Metadata Management
Standard Metadata Fields
Common fields stored in AudioTask data:
Character error rate (CER) is available as a utility function and typically requires a custom stage to compute and store it.
Error Handling
Graceful Failure Modes
AudioTask handles various error conditions:
Performance Characteristics
Memory Usage
AudioTask memory footprint is minimal since each task holds a single manifest entry. Memory scales with the number of metadata fields per entry and the total number of tasks processed in the pipeline.
Processing Patterns
Audio stages use several processing patterns:
Integration with Processing Stages
Stage Input/Output
AudioTask serves as input and output for most audio transform and filter stages, which subclass ProcessingStage[AudioTask, AudioTask]. Source and conversion stages can use different task types; for example, AudioToDocumentStage is ProcessingStage[AudioTask, DocumentBatch].
Chaining Stages
AudioTask flows through multiple processing stages, with each stage adding new metadata fields: