Migrate to NeMo Curator 26.09
Use this draft checklist to review applications and environments for the planned 26.09 release. The package version and publication status are not confirmed. For the feature summary, refer to the 26.09 release notes.
Migration at a Glance
Review the changes that affect your workflows:
Review Semantic Deduplication Precision
Pairwise compute defaults to float16 and processes ranked neighbors in batches of 1,024. Keep the default after validating duplicate outputs for your data. When your workflow requires float32 precision, set both kmeans_embedding_output_dtype="float32" and pairwise_compute_dtype="float32". Combining the default float16 KMeans output with float32 pairwise compute raises a ValueError at construction time. Validate results when changing precision.
KMeans can fit centroids from a sample of complete Parquet files and then reread all input for assignment. Review fit_data_fraction and leave GPU memory for reading, prediction, and writing. Refer to Semantic Deduplication for the current configuration.
Update Worker Configuration
Consider stage.with_(num_workers_per_node=...) when worker capacity should follow the number of live nodes. Do not combine it with num_workers or Ray Data actor-pool min/max/initial settings. Existing Xenna stage-spec values remain compatible when the common hook is unset. Refer to Stage Worker Sizing.
Check Tooling and Inference Dependencies
Use uv 0.12 or later to satisfy the project requirement. For GB200 vLLM startup, use quack-kernels>=0.4.1.
Use the installation command shown on the current installation page.
For container deployments, build a container image from the source release you adopt. The repository Dockerfile remains maintained. Audio and video workflows should use a derived image with the appropriate FFmpeg package or build; see Add FFmpeg for Audio or Video. NVIDIA no longer publishes new NGC images starting with 26.09, though previously published images remain downloadable. Refer to the Distribution Change.
The all extra no longer includes cv2 directly, though OpenCV may be installed transitively through vLLM. Include the cv2 extra for workflows that require it explicitly. Audio and video workflows should use a container image that includes the required FFmpeg system packages; Python package users can install FFmpeg on each executor as described in the Installation Guide.
Update Nemotron-Parse Examples
Replace unsupported writer imports or constructor arguments with InterleavedParquetWriterStage(path=...). In-process and Ray Serve inference select architecture-aware attention defaults. Preserve an explicit attention_backend only when your deployment requires it. Refer to Nemotron-Parse PDF processing.
Validate Before Adoption
Validate the updated environment and pipeline outputs before deployment:
- Confirm the final release tag, package version, and lockfile before upgrading.
- Reinstall the extras used by your pipeline with the release lockfile, or build an image from that release’s source.
- Run semantic deduplication on a representative sample and compare duplicate IDs at the selected precision.
- Run a representative multi-node pipeline and verify worker placement and output manifests.
- Validate Nemotron-Parse output on each GPU architecture used in production.