About NeMo CuratorRelease Notes

Overview

View as Markdown

Release Notes for NeMo Curator 26.09

NeMo Curator 26.09 expands audio inference and evaluation, adds Lance and remote WARC data sources, tightens semantic-deduplication memory use, and delivers pipeline scheduling improvements.

Before upgrading, review the 26.09 migration checklist for the OpenCV packaging change and the semantic-deduplication precision defaults.

Release 26.09

Release 26.09 corresponds to package version v1.4.0. It changes container distribution and adds the features, improvements, and fixes described below.

Distribution Change

Starting with 26.09, NVIDIA no longer publishes new NeMo Curator container images on NGC. Source releases continue on GitHub. Existing images remain available for download. No removal date has been announced.

For general workflows on 26.09 and later, install NeMo Curator using the appropriate method as described in the Installation Guide and Container Environments.

Highlights

New stages and data sources expand audio, text, and multimodal curation.

New ASR Adapters

Four new in-process automatic speech recognition (ASR) adapters expand the adapter surface for ASRStage:

  • QwenOmniASRAdapter: Qwen-Omni in-process transcription added in #1967.
  • Qwen3ASRAdapter: Qwen3-ASR in-process transcription added in #2257.
  • NeMoASRAdapter: NeMo FastConformer family transcription added in #2254.
  • FasterWhisperASRAdapter: Faster-Whisper in-process transcription added in #2271.

Refer to ASR Inference for adapter configuration and resource requirements.

Audio Transcript Processing Stages

Three new stages extend the audio transcript processing pipeline documented in Audio Text Integration:

  • SelectBestPredictionStage: Compares predictions from multiple ASR adapters and writes the highest-scoring prediction to a configurable output key (#2266).
  • WhisperHallucinationStage and TranscriptNormalizationStage: Rule-based hallucination filtering and text normalization for ASR output, including language-aware long-word thresholding (#2258, #2267).

Sound Event Detection

Two new stages add sound event detection to audio pipelines:

  • SoundEventDetectionStage: Runs inference to detect sound events within audio segments (#2260).
  • SoundEventPostprocessingStage: Applies postprocessing to raw detection outputs, including filtering and formatting (#2261).

LLM Judge Evaluation Suite

LLMJudgeWorkflow in nemo_curator.eval.llm_judge turns large language model (LLM) evaluation into a Curator data pipeline (#2324). You define judges, rubrics, and score filters in YAML and Jinja; the workflow starts a local Dynamo or vLLM inference server, runs NeMo Data Designer for structured judgment, and writes or filters results at dataset scale. Refer to LLM Judge Workflow and the LLM Judge Runner tutorial.

Lance Data Sources

Two new stages add support for the Lance columnar format:

  • LanceReaderStage: Reads Lance datasets (#2111).
  • LanceWriterStage: Writes Lance datasets (#2112).
  • InterleavedLanceReaderStage: Reads Lance datasets in interleaved mode for image-text and multimodal pipelines (#2232).

Remote WARC Support

Common Crawl WARC iterators now accept remote URLs in addition to local paths (#2300), enabling direct streaming from S3 or HTTPS sources without a local download step.

GlotLID Language and Script Identification

GlotLIDStage adds language and script identification using the GlotLID model, which supports over 1,600 language varieties (#2281). Use it alongside or as an alternative to FastTextLangId when finer language or script granularity is required.

Improvements

Updates to deduplication, inference, and scheduling improve memory use and pipeline control.

Bounded Semantic Deduplication Pairwise Processing

Pairwise similarity now uses bounded batch sizes and defaults to float16 compute (#2318, #2388):

  • pairwise_batch_size defaults to 1024, capping GPU memory consumption per batch.
  • pairwise_compute_dtype defaults to "float16", reducing memory use at the cost of reduced precision.

If you set pairwise_compute_dtype="float32", also set kmeans_embedding_output_dtype="float32". Refer to the 26.09 migration checklist for the required configuration change.

KMeans Fit Memory Reduction

KMeansStage reduces memory overhead during centroid fitting by reading Parquet files in bounded batches before full-dataset assignment (#2386). This extends the file-level sampling introduced in 26.07 to the full fit path.

Text Normalization in Deduplication

Exact and fuzzy deduplication workflows now support optional text normalization before signature generation (#2247). Normalization removes case and punctuation variation before hashing, so near-identical records that differ only in formatting are treated as duplicates.

Custom MinHash Metrics

MinhashStage exposes additional cardinality and similarity metrics (#2274). Pass extra_metrics to collect per-document and per-cluster statistics without a separate stage.

FastText Language ID Language Parameter

FastTextLangIdStage now accepts an optional language parameter (#2199). Set it to filter records to a target language code at the identification stage rather than adding a downstream filter.

Deduplication ID Field Control

Exact and fuzzy deduplication workflows accept drop_id_field=True to remove the deduplication ID column from the output when downstream stages do not require it (#2078).

Multiple Input Task Types Per Stage

ProcessingStage now allows its process method to declare more than one accepted input task type (#2239). This removes the need for wrapper stages that unify input types before a shared processing step.

Per-Node Stage Worker Sizing

Stages can opt into per-node worker counts with num_workers_per_node (#2354). When set, the executor scales workers proportionally to the number of nodes in the cluster rather than using a fixed count. Refer to Stage Worker Sizing.

Ray Data Scheduler Observability

Ray Data pipelines now emit scheduler diagnostics that report queue depth, stall causes, and resource wait times (#2293). Use these to distinguish underprovisioned resources from slow stages.

Configurable Resource Waits Between Actor Pool Stages

RayActorPoolExecutor stages now wait for required resources to become available before launching the next stage (#2357). Configure resource_wait_timeout_s to set the maximum wait before an error is raised.

Arrow Schema Preservation Through Parquet

Parquet writers now preserve the Arrow schema from the input task, including field metadata and nested type annotations (#2360). Previously, schemas could be inferred at write time, which dropped metadata.

Bounded vLLM Embedding Generation

VLLMEmbeddingStage now processes inputs in bounded batches (#2317), preventing out-of-memory conditions when embedding large partitions.

Nemotron-Parse Triton Attention

In-process and Ray Serve Nemotron-Parse inference now uses Triton-based attention by default on Ampere and Blackwell GPUs (#2393). This replaces the FlashAttention path and removes the flash-attn runtime requirement for PDF parsing. Refer to Nemotron-Parse PDF processing.

Fixes

This release fixes the following issues:

  • GB300 and large-memory GPU support: Resources(gpu_memory_gb=N) now floors the derived fractional GPU count at 0.1, preventing stages from resolving to 0 GPUs on devices with more than approximately 200 GB of total memory such as GB300 (#2421).
  • Semantic deduplication storage options: Parquet storage options passed to TextSemanticDeduplicationWorkflow are now forwarded to the output writer (#2316).
  • Semantic deduplication ID preservation: Semantic deduplication ID fields are preserved in the output when drop_id_field is not set (#2358).
  • Partial rerun routing: Partial pipeline reruns now route tasks to the current attempt instead of a prior completed attempt (#2381).
  • Nemotron OCR API key: The Nemotron OCR pipeline now uses a standardized API key field across in-process and remote serving configurations (#2405).
  • Empty QA batch handling: Diverse QA stages handle empty batches without raising an exception (#2450).
  • Interleaved batch string operations: Interleaved pipeline stages correctly apply pandas string operations across heterogeneous batch types (#2434).
  • Null WebDataset content types: Interleaved readers handle null or missing content-type fields in WebDataset shards without failing (#2438).
  • Hosted LLM tutorial default: Tutorial examples replace the retired meta/llama-3.3-70b-instruct default with nvidia/nemotron-3-super-120b-a12b across CLI defaults, notebooks, and README references (#2461).

Dependency Changes

Review these dependency changes before updating your environment:

  • OpenCV is now an optional extra: The all extra no longer includes cv2 directly. Install the cv2 extra for workflows that require OpenCV explicitly. vLLM can still install OpenCV as a transitive dependency. Refer to the 26.09 migration checklist and Install NeMo Curator.
  • datasets is a core dependency: huggingface_hub-based dataset access through the datasets library is now included in all installation extras (#2164).
  • RAPIDS 26.08: The 26.09 release uses RAPIDS 26.08 for GPU-accelerated text and semantic deduplication (#2303).
  • vLLM GB200 startup fix: The vLLM extra requires quack-kernels>=0.4.1 for GB200 startup compatibility (#2391).

Upgrade Checklist

Complete these checks before adopting 26.09:

  1. Apply every relevant code change in the 26.09 migration checklist.
  2. If your pipeline uses OpenCV features, add the cv2 extra: nemo-curator[text_cuda12,cv2].
  3. If you use pairwise_compute_dtype="float32", also set kmeans_embedding_output_dtype="float32" to avoid a ValueError at construction time.
  4. Reinstall the modality extras used by your pipeline and verify RAPIDS 26.08 compatibility.
  5. Run a representative pipeline with a small input set and inspect output schemas, deduplication IDs, and worker counts.