Get Started

Install (All Modalities)

View as Markdown

Installation Guide for NeMo Curator

Install NeMo Curator for text, image, video, and audio workflows, then verify the modules and GPU support your pipeline needs. For a single-modality path, start with a modality quickstart.

Prerequisites

Before installing, prepare a supported Python environment and the hardware required by your selected extras:

  1. Use Ubuntu 20.04, 22.04, or 24.04 for the recommended Linux setup.
  2. Install Python 3.11, 3.12, or 3.13.
  3. Provide at least 16 GB RAM for basic text processing; GPU acceleration is optional and benefits from an NVIDIA GPU with 16 GB or more VRAM.
  4. Install CUDA 12 and a compatible NVIDIA driver when using a GPU extra.

Python 3.10 support ended in NeMo Curator 26.07. 26.04 was the last release to support Python 3.10. Set up new environments on Python 3.11, 3.12, or 3.13. Refer to the 26.07 migration checklist for details.

Development and Production Environments

Use the deployment guides to plan resources for production or multi-node clusters:

Deployment Paths

Use CaseRequirementsGuide
Local DevelopmentMinimum specs listed aboveContinue below
Production ClustersDetailed hardware, network, storage specsDeployment Requirements
Multi-node SetupAdvanced infrastructure planningDeployment Options

Installation Methods

For most workflows, install NeMo Curator from PyPI when the package version you need is available. Use a source checkout for development or an unpublished release. Audio and video workflows are easiest to reproduce in a locally built container, where you can include the required FFmpeg system packages on every worker.

Starting with 26.09, NVIDIA no longer publishes new NeMo Curator container images on NGC. Existing images remain downloadable. GitHub source and releases continue.

Choose an installation method based on your environment:

Install a specific package version only after confirming it is available on PyPI and matches the release notes.

  1. Install uv and create an environment.

    curl -LsSf https://astral.sh/uv/0.12.1/install.sh | sh
    source $HOME/.local/bin/env
    uv venv
    source .venv/bin/activate
  2. Install NeMo Curator with all extras. The all extra includes text_cuda12, so this command uses Curator’s tested dependency override.

    curl -O https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/requirements/text_cuda12-overrides.txt
    uv pip install --torch-backend cu129 torch wheel_stub psutil setuptools setuptools_scm
    uv pip install --no-build-isolation \
    --override text_cuda12-overrides.txt \
    --torch-backend cu129 \
    --extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
    "nemo-curator[all]"

Source installs with uv sync --all-extras read the equivalent override directly from pyproject.toml.

Package Extras

Use installation extras to add only the components your pipeline needs. The following table lists the available extras and install commands.

Available Package Extras

ExtraInstallation CommandDescription
text_cpuuv pip install "nemo-curator[text_cpu]"CPU-only text processing and filtering
text_cuda12uv pip install --override text_cuda12-overrides.txt "nemo-curator[text_cuda12]"GPU-accelerated text processing with RAPIDS and vLLM; use the CUDA indexes shown below
audio_cpuuv pip install "nemo-curator[audio_cpu]"CPU-only audio curation with NeMo Toolkit ASR
audio_cuda12uv pip install "nemo-curator[audio_cuda12]"GPU-accelerated audio curation. uv installations require a transformers==4.55.2 override.
image_cpuuv pip install "nemo-curator[image_cpu]"CPU-only image processing
image_cuda12uv pip install "nemo-curator[image_cuda12]"GPU-accelerated image processing with NVIDIA DALI
video_cpuuv pip install "nemo-curator[video_cpu]"CPU-only video processing
video_cuda12uv pip install --no-build-isolation "nemo-curator[video_cuda12]"GPU-accelerated video processing with CUDA libraries. Requires FFmpeg and additional build dependencies when using uv.
cv2uv pip install "nemo-curator[cv2]"Optional OpenCV-based image and video operations
inference_serveruv pip install "nemo-curator[inference_server]"Ray Serve + vLLM for serving LLMs alongside curation pipelines
sdg_cpuuv pip install "nemo-curator[sdg_cpu]"CPU-only synthetic data generation with Data Designer
sdg_cuda12uv pip install "nemo-curator[sdg_cuda12]"GPU-accelerated synthetic data generation with local inference server support

Optional OpenCV Support

The base NeMo Curator installation does not install OpenCV. The cv2 extra installs opencv-python-headless for features that decode or transform images with OpenCV. Install it in the same environment as Curator when you use one of these features:

  • interleaved image decoding, InterleavedBlurFilterStage, InterleavedQRCodeFilterStage, or InterleavedCLIPScoreFilterStage;
  • Nemotron-Parse PDF postprocessing; or
  • video target-resolution frame resizing or motion-vector filtering.

For a PyPI installation, run:

uv pip install "nemo-curator[cv2]"

For a source checkout, run:

uv sync --extra cv2

The cv2 extra is intentionally separate from the all extra. This keeps OpenCV’s bundled media dependencies out of installations that do not use OpenCV-gated features. The extra does not install or configure the ffmpeg command-line tool. Video and audio workflows that call FFmpeg still require the FFmpeg setup described above.

Most OpenCV-gated features raise an ImportError that identifies the missing opencv-python-headless dependency and suggests installing nemo_curator[cv2]. Nemotron-Parse PDF processing instead logs a warning and skips the affected PDF when OpenCV is unavailable. To verify the active environment, run:

python -c "import cv2; print(cv2.__version__)"

When you install inference_server or an extra that includes it with pip, add the public PyTorch and vLLM CUDA 12.9 wheel indexes:

python -m pip install \
--extra-index-url https://download.pytorch.org/whl/cu129 \
--extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
"nemo-curator[inference_server]"

All remaining dependencies, including Dynamo, NIXL, and RAPIDS packages, are resolved from public PyPI.

For text_cuda12, download Curator’s override file and use uv pip install with the CUDA 12.9 wheel source:

curl -O https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/requirements/text_cuda12-overrides.txt
uv pip install \
--override text_cuda12-overrides.txt \
--torch-backend cu129 \
--extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
"nemo-curator[text_cuda12]"

For development tools such as pre-commit, ruff, and pytest, run uv sync --group dev --group linting --group test. The project manages these tools as dependency groups, not optional dependencies.

Standard pip is not supported for text_cuda12 or for installing all extras together. vLLM 0.22 requires numba==0.65.0, while RAPIDS 25.10 declares numba<0.62. Use uv pip install with the explicit override shown above or uv sync from a source checkout. Refer to Build an Image for container installation guidance. Plain uv pip install without the override cannot resolve this combination either.

Additional Setup

Install external tools that your selected stages require.

Install FFmpeg for Python Package Installs

Some Curator audio and video stages use the ffmpeg executable for resampling, decoding, encoding, or metadata extraction. Python extras do not install it. For audio and video workflows, the container guide shows how to package FFmpeg in an image used by every worker. For Python package installs, verify that ffmpeg is on PATH on every executor node, or install it using one of these options.

If you do not have root access, install an organization-approved FFmpeg package in a user-managed Conda or Mamba environment. For example:

conda install -c conda-forge ffmpeg
command -v ffmpeg
ffmpeg -version

Activate the same environment on every executor node before starting the pipeline. A standard FFmpeg build is sufficient for audio resampling. The video script tab below describes the maintained encoder-focused build.

The repository video build requires the CUDA toolkit (nvcc): If you encounter ERROR: failed checking for nvcc while running the video script, install the CUDA toolkit and make nvcc available on your PATH. Run nvcc --version to verify the installation.

Advanced Software Codec Support

Curator’s strict FFmpeg build routes H.264/HEVC/AV1 decode through NVDEC and excludes software H.264 encoders by default. You can add software codec support when CPU-only stages run ffprobe on H.264/HEVC/AV1 inputs or when you use --transcode-encoder=libopenh264.

Inside a container, run:

# Add software h264/hevc/av1 decoders only
bash /opt/Curator/docker/common/install_h264_support.sh
# Add those decoders plus the libopenh264 software h264 encoder
bash /opt/Curator/docker/common/install_h264_support.sh --with-libopenh264

With --with-libopenh264, the resulting FFmpeg binary links Cisco OpenH264. You are responsible for any license obligations imposed by the resulting binaries.


Installation Verification

After installation, run an import check. A successful check prints the installed version and confirms core modules can be imported:

  1. Import NeMo Curator and core pipeline modules.

    import nemo_curator
    print(f"NeMo Curator version: {nemo_curator.__version__}")
    from nemo_curator.pipeline import Pipeline
    from nemo_curator.tasks import DocumentBatch
    print("Core modules imported successfully")
  2. If you installed GPU support, check that the driver and GPU modules are available. A GPU-enabled environment prints the device and cuDF confirmation. A CPU-only environment can report that no GPU was detected.

    try:
    import torch
    if torch.cuda.is_available():
    print(f"GPU available: {torch.cuda.get_device_name(0)}")
    print(f"GPU memory: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")
    else:
    print("No GPU detected")
    import cudf
    print("cuDF available for GPU-accelerated deduplication")
    except ImportError as e:
    print(f"Some GPU modules not available: {e}")

Next Steps

Run a modality quickstart or review the deployment guidance for your target environment: