Curate AudioProcess DataQuality Assessment

Duration Filtering

View as Markdown

Filter audio samples by duration ranges, speech rate metrics, and temporal characteristics to create optimal datasets for ASR training and speech processing applications.

Duration-Based Quality Control

Why Duration Matters

Training Efficiency: Duration filtering can improve ASR training by removing samples that may be problematic for training

Processing Performance: Duration affects computational requirements:

  • Memory usage scales with audio length
  • Batch processing efficiency varies with duration variance
  • GPU utilization optimized for consistent lengths

Basic Duration Filtering

Simple Duration Range

from nemo_curator.stages.audio.common import GetAudioDurationStage, PreserveByValueStage
# Calculate duration for each audio file
duration_stage = GetAudioDurationStage(
audio_filepath_key="audio_filepath",
duration_key="duration"
)
# Filter for optimal duration range (1-15 seconds)
min_duration_filter = PreserveByValueStage(
input_value_key="duration",
target_value=1.0,
operator="ge" # greater than or equal
)
max_duration_filter = PreserveByValueStage(
input_value_key="duration",
target_value=15.0,
operator="le" # less than or equal
)
# Add to pipeline
pipeline.add_stage(duration_stage)
pipeline.add_stage(min_duration_filter)
pipeline.add_stage(max_duration_filter)

Use Case-Specific Ranges

# Duration ranges for different applications
duration_configs = {
"asr_training": {
"min_duration": 1.0,
"max_duration": 20.0,
"optimal_range": (2.0, 10.0)
},
"voice_cloning": {
"min_duration": 3.0,
"max_duration": 10.0,
"optimal_range": (4.0, 8.0)
},
"speech_synthesis": {
"min_duration": 2.0,
"max_duration": 15.0,
"optimal_range": (3.0, 12.0)
},
"keyword_spotting": {
"min_duration": 0.5,
"max_duration": 3.0,
"optimal_range": (1.0, 2.0)
}
}
def create_use_case_duration_filter(use_case: str) -> list[PreserveByValueStage]:
"""Create duration filters for specific use case."""
config = duration_configs.get(use_case, duration_configs["asr_training"])
return [
PreserveByValueStage(
input_value_key="duration",
target_value=config["min_duration"],
operator="ge"
),
PreserveByValueStage(
input_value_key="duration",
target_value=config["max_duration"],
operator="le"
)
]

Speech Rate Analysis

Speech rate metrics (words per second, characters per second) help identify samples with speaking speeds appropriate for your use case.

Calculate Speech Rate Metrics

ComputeWERStage calculates speech rates together with WER and CER. For top-level records, it requires a positive numeric duration. If the manifest does not already include that field, add GetAudioDurationStage first. Both configured text fields must exist, and the reference must normalize to nonempty text. An empty reference produces metrics.wer = None, metrics.cer = None, and metrics.metric_skip_reason = "empty_reference" without adding word_rate or char_rate. Then configure the hypothesis and reference fields used by your manifest:

from nemo_curator.stages.audio.common import GetAudioDurationStage
from nemo_curator.stages.audio.metrics.wer import ComputeWERStage
pipeline.add_stage(GetAudioDurationStage(
audio_filepath_key="audio_filepath",
duration_key="duration",
))
pipeline.add_stage(ComputeWERStage(
hypothesis_text_key="pred_text",
reference_text_key="text",
))
# Adds metrics.word_rate and metrics.char_rate when the prerequisites above are met.

For records that contain segments, each segment must instead have numeric start and end timestamps that define a positive interval. A missing or non-positive interval causes ComputeWERStage to emit 0.0 for both speech-rate metrics.

Speech Rate Filtering

If you have pre-calculated speech rate metrics in your data, you can filter based on them:

from nemo_curator.stages.audio.common import PreserveByValueStage
from nemo_curator.pipeline import Pipeline
# Example: Filter by speech rate if you have word_rate field in your data
pipeline = Pipeline(name="speech_rate_filtering")
# Filter by speech rate (1.5-5 words per second)
pipeline.add_stage(
PreserveByValueStage(
input_value_key="word_rate", # Assumes this field exists in your data
target_value=1.5,
operator="ge"
)
)
pipeline.add_stage(
PreserveByValueStage(
input_value_key="word_rate",
target_value=5.0,
operator="le"
)
)

PreserveByValueStage reads top-level fields. Before using the following filters, use a custom stage to copy metrics.word_rate and metrics.char_rate into top-level word_rate and char_rate fields. Alternatively, provide these fields in the input manifest.

Filtering by Speech Rate

After you calculate speech rate metrics, filter samples to keep those with appropriate speaking speeds:

Normal Speech Rate Range

from nemo_curator.stages.audio.common import PreserveByValueStage
# Filter by word rate (assumes word_rate field exists in your data)
word_rate_min_filter = PreserveByValueStage(
input_value_key="word_rate",
target_value=1.5,
operator="ge"
)
word_rate_max_filter = PreserveByValueStage(
input_value_key="word_rate",
target_value=5.0,
operator="le"
)
# Filter by character rate (assumes char_rate field exists in your data)
char_rate_min_filter = PreserveByValueStage(
input_value_key="char_rate",
target_value=8.0,
operator="ge"
)
char_rate_max_filter = PreserveByValueStage(
input_value_key="char_rate",
target_value=30.0,
operator="le"
)

These examples assume the nested values emitted by ComputeWERStage have been projected to the top-level word_rate and char_rate fields expected by PreserveByValueStage.

Normal Speech Rate Ranges

Typical speech rates for different contexts:

ContextWords/SecondCharacters/SecondUse Case
Slow/Clear Speech1.5 - 2.58 - 15Educational content, accessibility
Normal Conversation2.5 - 4.015 - 24General ASR training
Fast Speech4.0 - 5.024 - 30News, presentations
Very Fast>5.0>30May indicate errors or problematic samples

Best Practices

Duration Filtering Strategy

  1. Analyze First: Understand your dataset’s duration distribution
  2. Use Case Alignment: Align duration ranges with intended use
  3. Progressive Filtering: Apply duration filters before computationally expensive stages
  4. Quality Correlation: Consider correlation between duration and other quality metrics

Common Pitfalls

Over-Filtering: Removing too much data

# Check retention rates before applying filters
retention_rate = filtered_count / original_count
if retention_rate < 0.5: # Less than 50% retained
print("Warning: Very aggressive filtering - consider relaxing thresholds")

Under-Filtering: Keeping problematic samples that may negatively impact training or processing efficiency.

Composed Example

The following function shows how duration extraction composes with ASR and WER filtering. For the maintained runnable FLEURS entrypoint and complete export configuration, use tutorials/audio/fleurs/main.py with tutorials/audio/fleurs/pipeline.yaml.

from nemo_curator.pipeline import Pipeline
from nemo_curator.stages.audio.common import GetAudioDurationStage, PreserveByValueStage
from nemo_curator.stages.audio.datasets.fleurs.create_initial_manifest import CreateInitialManifestFleursStage
from nemo_curator.stages.audio.inference.asr.stage import ASRStage
from nemo_curator.stages.audio.metrics.wer import GetPairwiseWerStage
from nemo_curator.stages.audio.io.convert import AudioToDocumentStage
from nemo_curator.stages.resources import Resources
def create_audio_pipeline(raw_data_dir: str, wer_threshold: float = 75.0) -> Pipeline:
"""Build a FLEURS pipeline with duration and WER filtering."""
pipeline = Pipeline(name="audio_inference", description="Inference audio and filter by WER threshold.")
# Load FLEURS dataset
pipeline.add_stage(
CreateInitialManifestFleursStage(
lang="hy_am",
split="dev",
raw_data_dir=raw_data_dir,
).with_(batch_size=4)
)
# ASR inference
pipeline.add_stage(
ASRStage(
adapter_target="nemo_curator.models.asr.nemo_asr.NeMoASRAdapter",
model_id="nvidia/stt_hy_fastconformer_hybrid_large_pc",
max_audio_sec_per_actor=240.0,
max_inference_duration_s=120.0,
local_bucketing=True,
audio_filepath_key="audio_filepath",
adapter_kwargs={"use_cuda_graph_decoder": False},
).with_(resources=Resources(gpus=1.0))
)
# Calculate WER
pipeline.add_stage(
GetPairwiseWerStage(
text_key="text",
pred_text_key="pred_text",
wer_key="wer_pct",
)
)
# Calculate duration
pipeline.add_stage(
GetAudioDurationStage(
audio_filepath_key="audio_filepath",
duration_key="duration"
)
)
# Filter by WER threshold
pipeline.add_stage(
PreserveByValueStage(
input_value_key="wer_pct",
target_value=wer_threshold,
operator="le"
)
)
# Convert to document format
pipeline.add_stage(AudioToDocumentStage().with_(batch_size=1))
return pipeline

The supplied FLEURS YAML also configures the writer, backend, Ray lifecycle, CPU fallback, and model-specific decoder compatibility. Prefer that maintained configuration for an end-to-end run.