Curate TextProcess Data

Process Data for Text Curation

View as Markdown

Process text data that you loaded through NeMo Curator’s pipeline architecture.

NeMo Curator provides tools for processing text data as part of a model training pipeline. These tools help you analyze, transform, and filter datasets to create high-quality training input.

How It Works

NeMo Curator’s text processing capabilities are organized into six main categories:

  1. Language Management: Handle multilingual content, translation, and language-specific processing.
  2. Content Processing and Cleaning: Clean, normalize, and transform text content.
  3. Deduplication: Remove duplicate and near-duplicate documents efficiently.
  4. Quality Assessment and Filtering: Score and remove low-quality content using heuristics and machine learning classifiers.
  5. Specialized Processing: Apply domain-specific processing for code and advanced curation tasks.
  6. Interleaved Datasets: Read, write, and filter MINT-1T-style image-text datasets.

Each category provides specific implementations optimized for different curation needs. The result is a cleaned and filtered dataset ready for model training.


Language Management

Handle multilingual content, translation, and language-specific processing requirements.

Content Processing and Cleaning

Clean, normalize, and transform text content for high-quality training data.

Deduplication

Remove duplicate and near-duplicate documents efficiently from your text datasets. All deduplication methods support both identification (finding duplicates) and removal (filtering them out) workflows.

Quality Assessment and Filtering

Score and remove low-quality content using heuristics and machine learning classifiers.

Specialized Processing

Domain-specific processing for code and advanced curation tasks.

Interleaved Datasets

Read, write, and filter MINT-1T-style image-text interleaved datasets across WebDataset and Parquet formats.