GPU Processing Guide
This guide explains how to use GPU acceleration in NVIDIA NeMo Curator for faster text data processing.
Setting Up GPU Support
To use GPU acceleration, you’ll need:
- NVIDIA GPU with CUDA support
- RAPIDS libraries installed (cuDF, RMM)
- PyTorch with CUDA support for model inference
Example: GPU-Accelerated Text Classification
Example: GPU-Accelerated Fuzzy Deduplication
Tune GPU Writes for the Target Filesystem
NeMo Curator sets KVIKIO_AUTO_DIRECT_IO_WRITE=0 by default. This disables
KvikIO direct writes because they can reduce write throughput on distributed
filesystems such as Lustre. Enabling direct writes may improve throughput when
writing to a local SSD:
Set the variable before importing NeMo Curator or starting the job. For a distributed job, make sure every worker inherits the same value. Benchmark both values with the target filesystem and workload before changing the default.
This setting applies to GPU-backed writes performed through cuDF and KvikIO, primarily the intermediate and output writes in deduplication workflows. It does not affect general CPU-based file writes.
GPU-Accelerated Modules
NVIDIA NeMo Curator provides these GPU-accelerated modules:
Data Processing
- Exact deduplication: GPU-optimized processing for duplicate detection
- Fuzzy deduplication: GPU-accelerated MinHash computation for approximate duplicates
- Semantic deduplication: GPU embeddings and similarity calculations for content-based deduplication
Text Classification
- Domain classification: English and multilingual content categorization
- Quality classification: Content quality assessment using GPU-accelerated models
- Safety models: AEGIS and Instruction Data Guard for content safety evaluation
- Educational content: FineWeb models for educational value scoring
- Content type classification: Automatic content type detection
- Task and complexity classification: Instruction complexity assessment