For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
LogoLogoNeMo Curator
DocumentationAPI Reference
DocumentationAPI Reference
  • Home
    • Welcome
  • About NeMo Curator
    • Overview
    • Key Features
  • Get Started
    • Overview
    • Install (All Modalities)
    • Text Quickstart
    • Image Quickstart
    • Video Quickstart
    • Audio Quickstart
  • Curate Text
    • Overview
      • Overview
    • Save and Export
  • Curate Images
    • Overview
    • Save and Export
  • Curate Video
    • Overview
    • Load Data
    • Save and Export
  • Curate Audio
    • Overview
    • Save and Export
  • Setup & Deployment
    • Overview
  • Reference
    • Overview
    • Related Tools
  • Welcome
  • Overview
  • Key Features
  • Overview
  • Deduplication
  • Resource Allocation
  • Streaming
  • Auto-Balancing
  • Throughput
  • Overview
  • Loading
  • Acquisition
  • Processing
  • Curation Pipeline
  • Overview
  • Loading
  • Data Processing
  • Data Export
  • Overview
  • Architecture
  • Abstractions
  • Data Flow
  • Overview
  • Curation Pipeline
  • Audio Task
  • ASR Pipeline
  • Quality Metrics
  • Manifests and Ingest
  • ALM Pipeline
  • Text Integration
  • Overview
  • 26.07 Migration Checklist
  • Migration FAQ
  • Migration Guide
  • Overview
  • Install (All Modalities)
  • Text Quickstart
  • Image Quickstart
  • Video Quickstart
  • Audio Quickstart
  • Overview
  • Overview
  • Recipe
  • Stage Reference
  • Overview
  • ArXiv
  • Common Crawl
  • Custom Sources
  • Nemotron-Parse PDF Pipeline
  • Read Existing Data
  • Wikipedia
  • Overview
  • Overview
  • Add IDs
  • Text Cleaning
  • Overview
  • vLLM Embedder
  • Overview
  • Exact Deduplication
  • Fuzzy Deduplication
  • Semantic Deduplication
  • Overview
  • Language Detection
  • Stopwords
  • Translation
  • Overview
  • Classifier
  • Distributed Classifier
  • Heuristic Filtering
  • Overview
  • Code Processing
  • Overview
  • Interleaved IO
  • Interleaved Filters
  • Save and Export
  • Overview
  • LLM Client Setup
  • Inference Server
  • NeMo Data Designer
  • Multilingual Q&A
  • Overview
  • Tutorial
  • Overview
  • Task Reference
  • Overview
  • Overview
  • Beginner Tutorial
  • Deduplication Workflow
  • Overview
  • TAR Archives
  • Overview
  • Overview
  • CLIP Embedder
  • Overview
  • Aesthetic Filter
  • NSFW Filter
  • Save and Export
  • Overview
  • Overview
  • Beginner Tutorial
  • Split and Dedup
  • Overview
  • Add Custom Environment
  • Add Custom Code
  • Add Custom Model
  • Add Custom Stage
  • Load Data
  • Overview
  • Clipping
  • Transcoding
  • Filtering
  • Embeddings
  • Deduplication
  • Frame Extraction
  • Captions Preview
  • Save and Export
  • Overview
  • Overview
  • Beginner Tutorial
  • ALM Tutorial
  • Long-Form Audio Cutting for ALM Pretraining
  • ReadSpeech Tutorial
  • Overview
  • Custom Manifests
  • FLEURS Dataset
  • Local Files
  • Overview
  • Overview
  • NeMo ASR Models
  • Overview
  • WER Filtering
  • Duration Filtering
  • Overview
  • Preprocessing Stages
  • VAD Segmentation
  • Band Filter
  • UTMOS Filter
  • SIGMOS Filter
  • Speaker Separation
  • AudioDataFilterStage Composite
  • Overview
  • Duration Calculation
  • Format Validation
  • Overview
  • Data Builder
  • Overlap Filtering
  • Text Integration
  • Save and Export
  • Overview
  • Overview
  • Requirements
  • Multi-Node Ray on Slurm
  • SLURM Job Arrays
  • Overview
  • Overview
  • Overview
  • Memory Management
  • Monitoring
  • GPU Processing
  • Resumable Processing
  • Execution Backends
  • Stage Worker Sizing
  • Per-Stage Runtime Environments
  • Container Environments
  • Related Tools
On this page
  • End-to-End Recipes
  • Key Concepts for Tutorial Success
Curate TextTutorials

Text Curation Tutorials

||View as Markdown|

Hands-on tutorials for text curation workflows are available in the tutorials/text directory of the NeMo Curator GitHub repository.

End-to-End Recipes

Nemotron-CLIMB Data Curation

Discover a pretraining data mixture by embedding, clustering, pruning, tokenizing, training proxy models, and fitting a performance predictor.

Key Concepts for Tutorial Success

Before diving into the tutorials, familiarize yourself with these essential NeMo Curator concepts:

Pipeline Architecture

Core processing stages and pipeline concepts for text curation workflows data-structures distributed

Quality Assessment

Scoring and filtering techniques used in tutorials heuristics classifiers

Data Loading

Loading data from various sources common-crawl custom-data

Distributed Classification

GPU-accelerated classification concepts gpu scalable

Previous
Overview
Next
Recipe
NVIDIANVIDIA
Developer-friendly docs for your API
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2026, NVIDIA Corporation.