Overview
Release Notes for NeMo Curator 26.09
NeMo Curator 26.09 expands audio inference and evaluation, adds Lance and remote WARC data sources, tightens semantic-deduplication memory use, and delivers pipeline scheduling improvements.
Before upgrading, review the 26.09 migration checklist for the OpenCV packaging change and the semantic-deduplication precision defaults.
Release 26.09
Release 26.09 corresponds to package version v1.4.0. It changes container distribution and adds the features, improvements, and fixes described below.
Distribution Change
Starting with 26.09, NVIDIA no longer publishes new NeMo Curator container images on NGC. Source releases continue on GitHub. Existing images remain available for download. No removal date has been announced.
For general workflows on 26.09 and later, install NeMo Curator using the appropriate method as described in the Installation Guide and Container Environments.
Highlights
New stages and data sources expand audio, text, and multimodal curation.
New ASR Adapters
Four new in-process automatic speech recognition (ASR) adapters expand the adapter surface for ASRStage:
QwenOmniASRAdapter: Qwen-Omni in-process transcription added in #1967.Qwen3ASRAdapter: Qwen3-ASR in-process transcription added in #2257.NeMoASRAdapter: NeMo FastConformer family transcription added in #2254.FasterWhisperASRAdapter: Faster-Whisper in-process transcription added in #2271.
Refer to ASR Inference for adapter configuration and resource requirements.
Audio Transcript Processing Stages
Three new stages extend the audio transcript processing pipeline documented in Audio Text Integration:
SelectBestPredictionStage: Compares predictions from multiple ASR adapters and writes the highest-scoring prediction to a configurable output key (#2266).WhisperHallucinationStageandTranscriptNormalizationStage: Rule-based hallucination filtering and text normalization for ASR output, including language-aware long-word thresholding (#2258, #2267).
Sound Event Detection
Two new stages add sound event detection to audio pipelines:
SoundEventDetectionStage: Runs inference to detect sound events within audio segments (#2260).SoundEventPostprocessingStage: Applies postprocessing to raw detection outputs, including filtering and formatting (#2261).
LLM Judge Evaluation Suite
LLMJudgeWorkflow in nemo_curator.eval.llm_judge turns large language model (LLM) evaluation into a Curator data pipeline (#2324). You define judges, rubrics, and score filters in YAML and Jinja; the workflow starts a local Dynamo or vLLM inference server, runs NeMo Data Designer for structured judgment, and writes or filters results at dataset scale. Refer to LLM Judge Workflow and the LLM Judge Runner tutorial.
Lance Data Sources
Two new stages add support for the Lance columnar format:
LanceReaderStage: Reads Lance datasets (#2111).LanceWriterStage: Writes Lance datasets (#2112).InterleavedLanceReaderStage: Reads Lance datasets in interleaved mode for image-text and multimodal pipelines (#2232).
Remote WARC Support
Common Crawl WARC iterators now accept remote URLs in addition to local paths (#2300), enabling direct streaming from S3 or HTTPS sources without a local download step.
GlotLID Language and Script Identification
GlotLIDStage adds language and script identification using the GlotLID model, which supports over 1,600 language varieties (#2281). Use it alongside or as an alternative to FastTextLangId when finer language or script granularity is required.
Improvements
Updates to deduplication, inference, and scheduling improve memory use and pipeline control.
Bounded Semantic Deduplication Pairwise Processing
Pairwise similarity now uses bounded batch sizes and defaults to float16 compute (#2318, #2388):
pairwise_batch_sizedefaults to1024, capping GPU memory consumption per batch.pairwise_compute_dtypedefaults to"float16", reducing memory use at the cost of reduced precision.
If you set pairwise_compute_dtype="float32", also set kmeans_embedding_output_dtype="float32". Refer to the 26.09 migration checklist for the required configuration change.
KMeans Fit Memory Reduction
KMeansStage reduces memory overhead during centroid fitting by reading Parquet files in bounded batches before full-dataset assignment (#2386). This extends the file-level sampling introduced in 26.07 to the full fit path.
Text Normalization in Deduplication
Exact and fuzzy deduplication workflows now support optional text normalization before signature generation (#2247). Normalization removes case and punctuation variation before hashing, so near-identical records that differ only in formatting are treated as duplicates.
Custom MinHash Metrics
MinhashStage exposes additional cardinality and similarity metrics (#2274). Pass extra_metrics to collect per-document and per-cluster statistics without a separate stage.
FastText Language ID Language Parameter
FastTextLangIdStage now accepts an optional language parameter (#2199). Set it to filter records to a target language code at the identification stage rather than adding a downstream filter.
Deduplication ID Field Control
Exact and fuzzy deduplication workflows accept drop_id_field=True to remove the deduplication ID column from the output when downstream stages do not require it (#2078).
Multiple Input Task Types Per Stage
ProcessingStage now allows its process method to declare more than one accepted input task type (#2239). This removes the need for wrapper stages that unify input types before a shared processing step.
Per-Node Stage Worker Sizing
Stages can opt into per-node worker counts with num_workers_per_node (#2354). When set, the executor scales workers proportionally to the number of nodes in the cluster rather than using a fixed count. Refer to Stage Worker Sizing.
Ray Data Scheduler Observability
Ray Data pipelines now emit scheduler diagnostics that report queue depth, stall causes, and resource wait times (#2293). Use these to distinguish underprovisioned resources from slow stages.
Configurable Resource Waits Between Actor Pool Stages
RayActorPoolExecutor stages now wait for required resources to become available before launching the next stage (#2357). Configure resource_wait_timeout_s to set the maximum wait before an error is raised.
Arrow Schema Preservation Through Parquet
Parquet writers now preserve the Arrow schema from the input task, including field metadata and nested type annotations (#2360). Previously, schemas could be inferred at write time, which dropped metadata.
Bounded vLLM Embedding Generation
VLLMEmbeddingStage now processes inputs in bounded batches (#2317), preventing out-of-memory conditions when embedding large partitions.
Nemotron-Parse Triton Attention
In-process and Ray Serve Nemotron-Parse inference now uses Triton-based attention by default on Ampere and Blackwell GPUs (#2393). This replaces the FlashAttention path and removes the flash-attn runtime requirement for PDF parsing. Refer to Nemotron-Parse PDF processing.
Fixes
This release fixes the following issues:
- GB300 and large-memory GPU support:
Resources(gpu_memory_gb=N)now floors the derived fractional GPU count at0.1, preventing stages from resolving to0GPUs on devices with more than approximately 200 GB of total memory such as GB300 (#2421). - Semantic deduplication storage options: Parquet storage options passed to
TextSemanticDeduplicationWorkfloware now forwarded to the output writer (#2316). - Semantic deduplication ID preservation: Semantic deduplication ID fields are preserved in the output when
drop_id_fieldis not set (#2358). - Partial rerun routing: Partial pipeline reruns now route tasks to the current attempt instead of a prior completed attempt (#2381).
- Nemotron OCR API key: The Nemotron OCR pipeline now uses a standardized API key field across in-process and remote serving configurations (#2405).
- Empty QA batch handling: Diverse QA stages handle empty batches without raising an exception (#2450).
- Interleaved batch string operations: Interleaved pipeline stages correctly apply pandas string operations across heterogeneous batch types (#2434).
- Null WebDataset content types: Interleaved readers handle null or missing content-type fields in WebDataset shards without failing (#2438).
- Hosted LLM tutorial default: Tutorial examples replace the retired
meta/llama-3.3-70b-instructdefault withnvidia/nemotron-3-super-120b-a12bacross CLI defaults, notebooks, and README references (#2461).
Dependency Changes
Review these dependency changes before updating your environment:
- OpenCV is now an optional extra: The
allextra no longer includescv2directly. Install thecv2extra for workflows that require OpenCV explicitly. vLLM can still install OpenCV as a transitive dependency. Refer to the 26.09 migration checklist and Install NeMo Curator. datasetsis a core dependency:huggingface_hub-based dataset access through thedatasetslibrary is now included in all installation extras (#2164).- RAPIDS 26.08: The 26.09 release uses RAPIDS 26.08 for GPU-accelerated text and semantic deduplication (#2303).
- vLLM GB200 startup fix: The vLLM extra requires
quack-kernels>=0.4.1for GB200 startup compatibility (#2391).
Upgrade Checklist
Complete these checks before adopting 26.09:
- Apply every relevant code change in the 26.09 migration checklist.
- If your pipeline uses OpenCV features, add the
cv2extra:nemo-curator[text_cuda12,cv2]. - If you use
pairwise_compute_dtype="float32", also setkmeans_embedding_output_dtype="float32"to avoid aValueErrorat construction time. - Reinstall the modality extras used by your pipeline and verify RAPIDS 26.08 compatibility.
- Run a representative pipeline with a small input set and inspect output schemas, deduplication IDs, and worker counts.