Process Data for Text Curation
Process text data you’ve loaded through NeMo Curator’s pipeline architecture .
NeMo Curator provides a comprehensive suite of tools for processing text data as part of the AI training pipeline. These tools help you analyze, transform, and filter your text datasets to ensure high-quality input for language model training.
How it Works
NeMo Curator’s text processing capabilities are organized into six main categories:
- Language Management: Handle multilingual content, translation, and language-specific processing
- Content Processing & Cleaning: Clean, normalize, and transform text content
- Deduplication: Remove duplicate and near-duplicate documents efficiently
- Quality Assessment & Filtering: Score and remove low-quality content using heuristics and ML classifiers
- Specialized Processing: Domain-specific processing for code and advanced curation tasks
- Interleaved Datasets: Read, write, and filter MINT-1T-style image-text datasets
Each category provides specific implementations optimized for different curation needs. The result is a cleaned and filtered dataset ready for model training.
Language Management
Handle multilingual content, translation, and language-specific processing requirements.
Content Processing & Cleaning
Clean, normalize, and transform text content for high-quality training data.
Deduplication
Remove duplicate and near-duplicate documents efficiently from your text datasets. All deduplication methods support both identification (finding duplicates) and removal (filtering them out) workflows.
Quality Assessment & Filtering
Score and remove low-quality content using heuristics and ML classifiers.
Specialized Processing
Domain-specific processing for code and advanced curation tasks.
Interleaved Datasets
Read, write, and filter MINT-1T-style image-text interleaved datasets across WebDataset and Parquet formats.