Text Integration for Audio Data
Convert processed audio data from AudioTask to DocumentBatch format using the built-in AudioToDocumentStage. This enables you to export audio processing results or integrate with custom text processing workflows.
How it Works
The AudioToDocumentStage provides straightforward format conversion between NeMo Curator’s audio and text data structures:
- Format Conversion: Transform
AudioTaskobjects toDocumentBatchformat - Metadata Sanitization: Preserve serializable top-level metadata while removing raw audio, segments, and tensor values
- Export Ready: Convert audio processing results to pandas DataFrame format for analysis or export
Common use cases:
- Export ASR results and quality metrics for analysis
- Save filtered audio datasets with transcriptions
- Integrate audio processing outputs with downstream text workflows
Basic Conversion
AudioTask to DocumentBatch
Use AudioToDocumentStage to convert audio processing results to document format:
Parameters:
AudioToDocumentStage()has no configuration parameters; it performs direct format conversion
Returns:
- A list containing one
DocumentBatchwhose pandas DataFrame contains the retained fields from the input batch
Preserved and Removed Fields
The conversion preserves ordinary top-level scalar, list, and dictionary fields, including:
For retained fields, names and values are preserved as they appear in the AudioTask. Before building the DataFrame, AudioToDocumentStage removes waveform, audio, audio_data, audio_array, segments, and every top-level tensor-valued field.
Preserve Segments During JSONL Export
If you need segments or segment-level WER/CER in the output, keep the records as AudioTask objects and use ManifestWriterStage instead of converting them:
ManifestWriterStage serializes task.data directly. Remove waveform arrays and tensor-valued fields before this writer because they are not JSON-serializable.
Process Transcripts Before Conversion
Use the audio transcript stages for cleanup and screening before you convert records to DocumentBatch. Add them to an existing audio pipeline that reads AudioTask records.
WhisperHallucinationStage is a lexical heuristic. A flag identifies a transcript for review or filtering. It does not prove that the transcript is factually wrong, and the stage does not drop the record. Existing non-hallucination skip reasons remain unchanged. With overwrite=True, a later clean transcript can clear an existing hallucination flag and write a recovery note. This re-check does not invoke ASR. Use a reference transcript or a second model when your workflow needs recovery.
For this example, each manifest record includes pred_text, duration, _skipme, and language. language selects the language-aware long-word rule: agglutinative and compounding languages use a 35-character threshold without the relative-length check. For languages written without word-separating spaces, the repeated-word and long-word checks are disabled; phrase matching and character-rate checks remain active. Set the correct language field so the stage can apply these rules. A missing language is treated as empty and does not activate the no-space-script handling. AbbreviationConcatStage defaults to source_lang and English. Set source_lang_key="language" to share the manifest language field. Refer to the maintained hallucination pipeline configuration and phrase corpus for the complete inputs.
Place transcript cleanup before GetPairwiseWerStage so the stored WER reflects the cleaned prediction. If cleanup runs afterward, recalculate WER.
Add transcript stages to an existing pipeline as follows. Use absolute paths for the input YAML and hallucination phrase corpus:
The regex file is a YAML list. Rules run in file order:
For records whose ASR text is stored in pred_text, configure the selector input explicitly:
Integration in Pipelines
Complete Audio Processing with Export
The most common use case is adding AudioToDocumentStage at the end of your audio pipeline to enable result export:
Output format: The JsonlWriter creates a JSONL file where each line contains one audio sample with the fields retained by AudioToDocumentStage:
Custom Integration
While AudioToDocumentStage converts audio data to DocumentBatch format, NeMo Curator’s built-in text processing stages (filters, classifiers, and so on) are designed for text documents, not audio transcriptions. For audio-specific text processing, implement custom stages that operate on the converted DocumentBatch data.
Example: Custom Text Processing
Output Format
After conversion, your data will be in DocumentBatch format with a pandas DataFrame:
Limitations
Text Processing Integration: NeMo Curator’s text processing stages are designed for DocumentBatch inputs (text documents such as articles, web pages), but they are not designed for audio-derived transcriptions. You should implement custom processing stages for audio-specific workflows.
Reasons for incompatibility:
- Text filters assume document-level content (e.g., paragraph structure, word count thresholds designed for articles)
- ASR transcriptions have different characteristics (shorter, can contain recognition errors, conversational language)
- Audio-specific metrics (WER, duration, speech rate) require custom filtering logic
Recommendation: Use PreserveByValueStage for audio quality filtering, or create custom stages for transcription-specific processing.
Related Topics
- Audio Processing Overview - Complete audio processing workflow
- Quality Assessment - Audio quality metrics and filtering
- ASR Inference - Speech recognition processing