Get Started

Quickstart for NeMo Curator Image Curation

View as Markdown

Use this quickstart to install NeMo Curator, run an image curation example, and inspect the output. The example reads JPEG images from tar archives and writes curated image shards with Parquet metadata.

Prerequisites

Before you start, prepare the following environment and input data:

  • Python 3.11, 3.12, or 3.13, with packaging >= 22.0 and uv.
  • Ubuntu 20.04 or 22.04.
  • An NVIDIA GPU with compute capability 7.0 or higher and CUDA 12 or later. Image modules require a GPU.
  • JPEG images in .tar archives. CLIP and classifier model weights download automatically on first use.
  • The image reader ignores text and JSON files in the archives.

Install uv if it is not already available. Refer to the NeMo Curator Installation Guide for setup details.

curl -LsSf https://astral.sh/uv/0.12.1/install.sh | sh
source $HOME/.local/bin/env

Quickstart Steps

Complete these four steps to install NeMo Curator, run the example, and verify the generated files.

  1. Install NeMo Curator with an option that matches your environment.

    This installs the development branch. To install a release, check out its published GitHub release tag before running uv sync.

    git clone https://github.com/NVIDIA-NeMo/Curator.git
    cd Curator
    uv sync --extra image_cuda12
    source .venv/bin/activate
  1. Create directories for the input archives, curated output, and model weights.

    mkdir -p ~/nemo_curator/data/tar_archives
    mkdir -p ~/nemo_curator/data/curated
    mkdir -p ~/nemo_curator/models

    Place JPEG images in .tar archives under ~/nemo_curator/data/tar_archives.

  2. Download and run the image curation example.

    wget -O ~/nemo_curator/image_curation_example.py https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/tutorials/image/getting-started/image_curation_example.py
    python ~/nemo_curator/image_curation_example.py \
    --input-wds-dataset-dir ~/nemo_curator/data/tar_archives \
    --output-dataset-dir ~/nemo_curator/data/curated \
    --model-dir ~/nemo_curator/models \
    --aesthetic-threshold 0.5 \
    --nsfw-threshold 0.5

    The Image Curation Tutorial contains the complete example.

  3. Confirm that the output directory contains curated tar files and matching Parquet metadata files. Refer to Expected Output for the file layout.

Basic Image Curation Example

This code builds an image curation pipeline with file partitioning, image loading, embeddings, quality filters, and output writing.

CPU Memory Considerations

Image loading and decoding use CPU memory before GPU processing. If ImageReaderStage runs out of memory, reduce one or more of these settings:

  • dali_batch_size: Reduce the number of images per batch to 32–50 on systems with limited RAM.
  • num_threads: Reduce parallel decoding threads to four on systems with limited RAM.
  • num_cpus: Reduce Ray Client CPU allocation to 8–16 on systems with limited RAM.

The code example uses conservative defaults. Increase these values on systems with more memory.

Use this code to limit the CPUs that Ray can allocate:

from nemo_curator.core.client import RayClient
ray_client = RayClient(num_cpus=8) # Adjust based on available CPU cores
ray_client.start()
from nemo_curator.pipeline import Pipeline
from nemo_curator.backends.xenna import XennaExecutor
from nemo_curator.stages.file_partitioning import FilePartitioningStage
from nemo_curator.stages.image.io.image_reader import ImageReaderStage
from nemo_curator.stages.image.embedders.clip_embedder import ImageEmbeddingStage
from nemo_curator.stages.image.filters.aesthetic_filter import ImageAestheticFilterStage
from nemo_curator.stages.image.filters.nsfw_filter import ImageNSFWFilterStage
from nemo_curator.stages.image.io.image_writer import ImageWriterStage
# Create image curation pipeline
pipeline = Pipeline(name="image_curation", description="Basic image curation with quality filtering")
# Stage 1: Partition tar files for parallel processing
pipeline.add_stage(FilePartitioningStage(
file_paths="~/nemo_curator/data/tar_archives", # Path to your tar archive directory
files_per_partition=1,
file_extensions=[".tar"],
))
# Stage 2: Read images from tar files using DALI
pipeline.add_stage(ImageReaderStage(
dali_batch_size=50,
verbose=True,
num_threads=4,
num_gpus_per_worker=0.25,
))
# Stage 3: Generate CLIP embeddings for images
pipeline.add_stage(ImageEmbeddingStage(
model_dir="~/nemo_curator/models", # Directory containing model weights
model_inference_batch_size=32,
num_gpus_per_worker=0.25,
remove_image_data=False,
verbose=True,
))
# Stage 4: Filter by aesthetic quality (keep images with score >= 0.5)
pipeline.add_stage(ImageAestheticFilterStage(
model_dir="~/nemo_curator/models",
score_threshold=0.5,
model_inference_batch_size=32,
num_gpus_per_worker=0.25,
verbose=True,
))
# Stage 5: Filter NSFW content (remove images with score >= 0.5)
pipeline.add_stage(ImageNSFWFilterStage(
model_dir="~/nemo_curator/models",
score_threshold=0.5,
model_inference_batch_size=32,
num_gpus_per_worker=0.25,
verbose=True,
))
# Stage 6: Save curated images to new tar archives
pipeline.add_stage(ImageWriterStage(
output_dir="~/nemo_curator/data/curated",
images_per_tar=1000,
remove_image_data=True,
verbose=True,
))
# Execute the pipeline
executor = XennaExecutor()
pipeline.run(executor)

Expected Output

After the pipeline runs, your output directory contains files like these:

~/nemo_curator/data/curated/
├── images-{hash}-000000.tar # Curated images (first shard)
├── images-{hash}-000000.parquet # Metadata for corresponding tar
├── images-{hash}-000001.tar # Curated images (second shard)
├── images-{hash}-000001.parquet # Metadata for corresponding tar
├── ... # Additional shards as needed

The output files contain the following data:

  • Tar Files: High-quality .jpg files that passed aesthetic and NSFW filtering.
  • Parquet Files: Metadata for each tar file, including image paths, IDs, and processing scores.
  • Naming Convention: Hash-based prefixes, such as images-a1b2c3d4e5f6-000000.tar, keep distributed outputs unique.
  • Scores: The Parquet files store aesthetic_score and nsfw_score values.

Complete Tutorial

The command in Quickstart Step 3 downloads and runs the complete example, including its data and configuration options.

Next Steps

Continue with these image curation references and the full tutorial: