Quality Assessment and Filtering
Score and remove low-quality content using heuristics and machine learning classifiers to prepare your data for model training.
Large datasets often contain many documents considered “low quality.” In this context, low-quality data is unsuitable for downstream model training. High-quality data contains the characteristics that you want downstream models to learn. The metrics that define quality vary widely.
How It Works
NeMo Curator’s filtering framework includes several components that work within the data processing architecture:
ScoreFilter
Filter and Score Modules
The ScoreFilter is at the center of filtering in NeMo Curator. It applies a filter to a document and optionally saves the score as metadata:
Default Executor: When you call pipeline.run() without specifying an executor, NeMo Curator automatically uses XennaExecutor() as the default. You can optionally specify a different executor by passing it as a parameter: pipeline.run(executor=my_executor).
The filter object implements two key methods:
score_document: Computes a quality score for a document.keep_document: Determines whether to keep a document based on its score.
Filtering Approaches
Usage
NeMo Curator provides programmatic interfaces for document filtering through the Pipeline framework:
Best Practices
When filtering large datasets, consider these performance tips:
- Order matters: Place computationally inexpensive filters early in your pipeline.
- Batch size tuning: Adjust batch sizes based on your hardware capabilities.
- Use vectorization: Implement batched methods for compute-intensive filters.
- Disk I/O: Consider compression and chunking strategies for large datasets.
- Distributed processing: For TB-scale datasets, use distributed filtering with the XennaExecutor.
Related Topics
Refer to the following related documentation: