Heuristic Filtering
Heuristic filtering uses simple, rule-based metrics to identify and remove low-quality documents from your dataset. NVIDIA NeMo Curator provides prebuilt heuristic filters that you can configure and combine for your requirements.
How It Works
Heuristic filters examine specific attributes of text documents and apply predefined thresholds to determine document quality. Unlike classifier-based filtering, heuristic filters do not require training data. They rely on configurable thresholds and rules.
These filters assess quality using measurable document characteristics such as:
- Document length (word or character count)
- Punctuation ratios and patterns
- Repetitive content detection
- Language-specific patterns
- Text completeness and coherence
For details on filter structure and the filtering process, refer to Data Processing Concepts .
Usage
Python
Configuration
Available Filters
NeMo Curator includes more than 30 heuristic filters for assessing document quality. Below are the most commonly used filters with their parameters:
Text Length Filters
The following filters measure document and word length:
Repetition Detection Filters
The following filters detect repeated lines, paragraphs, or n-grams:
Character and Symbol Filters
The following filters measure character, symbol, numeric, URL, punctuation, and whitespace patterns:
Content-Specific Filters
The following filters identify content-specific quality signals:
Special Purpose Filters
The following filters address specialized quality conditions:
Configuration
NeMo Curator pipelines can be configured using YAML files with Hydra. The configuration uses _target_ to specify class paths:
Hydra Configuration
Refer to nemo_curator/config/text/ for complete pipeline examples.
For non-English text, you can adjust the filter parameters based on the characteristics of your target language.
Best Practices
When building filter chains, follow these best practices:
Order for Efficiency
Performance Tuning
Precision vs. Recall
Language Considerations
Multiple Filters
Analyzing Filter Results
When tuning filter thresholds, analyze score distributions before applying filters. NeMo Curator provides two modules for this workflow:
Score: Computes scores and adds them as columns without removing documents.ScoreFilter: Computes scores, filters based on thresholds, and optionally retains scores in output.
Use Score first to understand your data distribution, then apply ScoreFilter with tuned thresholds.
By default, Filter and ScoreFilter log when an input batch is already empty or when a stage removes every row, but they do not report partial retention. Set verbose=True to also log <stage> batch <task_id> retained <kept>/<input> rows for each non-empty batch. A stage containing multiple filters reports one count after the complete chain. These messages contain aggregate counts, not the identities or rejection reasons of individual records.
Score Without Filtering
Analyze Score Distribution
Apply Tuned Filters
Use Score to add score columns to your data without removing any documents:
Output files are written to the scored_output/ directory with one file per input partition.
Performance Tuning
For large datasets, consider these performance optimizations:
XennaExecutor (Default)
RayDataExecutor
Batch Size Optimization
XennaExecutor is the default executor, optimized for streaming workloads. You can customize its configuration or use the defaults:
If no executor is specified, pipeline.run() uses XennaExecutor with default settings.
Pipeline Metrics
When you run a filtering pipeline, each stage tracks the number of documents it processes. You can use these metrics to understand how each filter affects your dataset and to tune thresholds.
After calling pipeline.run(), the returned task objects contain per-stage performance statistics through _stage_perf. Each entry is a StagePerfStats object with a num_items_processed field that records how many documents passed through that stage.
These same metrics power the nightly benchmarks, which track num_documents_processed, num_kept_documents, and throughput_docs_per_sec for every pipeline run.
Remember that the goal of filtering is to improve the quality of your training data, not necessarily to remove as many documents as possible. Monitor your filtering results and adjust thresholds based on your specific data characteristics and downstream tasks.