Duration Filtering
Filter audio samples by duration ranges, speech rate metrics, and temporal characteristics to create optimal datasets for ASR training and speech processing applications.
Duration-Based Quality Control
Why Duration Matters
Training Efficiency: Duration filtering can improve ASR training by removing samples that may be problematic for training
Processing Performance: Duration affects computational requirements:
- Memory usage scales with audio length
- Batch processing efficiency varies with duration variance
- GPU utilization optimized for consistent lengths
Basic Duration Filtering
Simple Duration Range
Use Case-Specific Ranges
Speech Rate Analysis
Speech rate metrics (words per second, characters per second) help identify samples with speaking speeds appropriate for your use case.
Calculate Speech Rate Metrics
ComputeWERStage calculates speech rates together with WER and CER. For top-level records, it requires a positive numeric duration. If the manifest does not already include that field, add GetAudioDurationStage first. Both configured text fields must exist, and the reference must normalize to nonempty text. An empty reference produces metrics.wer = None, metrics.cer = None, and metrics.metric_skip_reason = "empty_reference" without adding word_rate or char_rate. Then configure the hypothesis and reference fields used by your manifest:
For records that contain segments, each segment must instead have numeric start and end timestamps that define a positive interval. A missing or non-positive interval causes ComputeWERStage to emit 0.0 for both speech-rate metrics.
Speech Rate Filtering
If you have pre-calculated speech rate metrics in your data, you can filter based on them:
PreserveByValueStage reads top-level fields. Before using the following filters, use a custom stage to copy metrics.word_rate and metrics.char_rate into top-level word_rate and char_rate fields. Alternatively, provide these fields in the input manifest.
Filtering by Speech Rate
After you calculate speech rate metrics, filter samples to keep those with appropriate speaking speeds:
Normal Speech Rate Range
These examples assume the nested values emitted by ComputeWERStage have been projected to the top-level word_rate and char_rate fields expected by PreserveByValueStage.
Normal Speech Rate Ranges
Typical speech rates for different contexts:
Best Practices
Duration Filtering Strategy
- Analyze First: Understand your dataset’s duration distribution
- Use Case Alignment: Align duration ranges with intended use
- Progressive Filtering: Apply duration filters before computationally expensive stages
- Quality Correlation: Consider correlation between duration and other quality metrics
Common Pitfalls
Over-Filtering: Removing too much data
Under-Filtering: Keeping problematic samples that may negatively impact training or processing efficiency.
Composed Example
The following function shows how duration extraction composes with ASR and WER filtering. For the maintained runnable FLEURS entrypoint and complete export configuration, use tutorials/audio/fleurs/main.py with tutorials/audio/fleurs/pipeline.yaml.
The supplied FLEURS YAML also configures the writer, backend, Ray lifecycle, CPU fallback, and model-specific decoder compatibility. Prefer that maintained configuration for an end-to-end run.
Related Topics
- Quality Assessment Overview: Complete quality filtering workflow
- WER Filtering: Transcription accuracy filtering
- Audio Analysis: Duration calculation and analysis