Quickstart for NeMo Curator Image Curation
Use this quickstart to install NeMo Curator, run an image curation example, and inspect the output. The example reads JPEG images from tar archives and writes curated image shards with Parquet metadata.
Prerequisites
Before you start, prepare the following environment and input data:
- Python 3.11, 3.12, or 3.13, with
packaging >= 22.0anduv. - Ubuntu 20.04 or 22.04.
- An NVIDIA GPU with compute capability 7.0 or higher and CUDA 12 or later. Image modules require a GPU.
- JPEG images in
.tararchives. CLIP and classifier model weights download automatically on first use. - The image reader ignores text and JSON files in the archives.
Install uv if it is not already available. Refer to the NeMo Curator Installation Guide for setup details.
Quickstart Steps
Complete these four steps to install NeMo Curator, run the example, and verify the generated files.
-
Install NeMo Curator with an option that matches your environment.
Source Installation
PyPI Installation
Build an Image
This installs the development branch. To install a release, check out its published GitHub release tag before running
uv sync.
-
Create directories for the input archives, curated output, and model weights.
Place JPEG images in
.tararchives under~/nemo_curator/data/tar_archives. -
Download and run the image curation example.
The Image Curation Tutorial contains the complete example.
-
Confirm that the output directory contains curated tar files and matching Parquet metadata files. Refer to Expected Output for the file layout.
Basic Image Curation Example
This code builds an image curation pipeline with file partitioning, image loading, embeddings, quality filters, and output writing.
CPU Memory Considerations
Image loading and decoding use CPU memory before GPU processing. If ImageReaderStage runs out of memory, reduce one or more of these settings:
dali_batch_size: Reduce the number of images per batch to 32–50 on systems with limited RAM.num_threads: Reduce parallel decoding threads to four on systems with limited RAM.num_cpus: Reduce Ray Client CPU allocation to 8–16 on systems with limited RAM.
The code example uses conservative defaults. Increase these values on systems with more memory.
Use this code to limit the CPUs that Ray can allocate:
Expected Output
After the pipeline runs, your output directory contains files like these:
The output files contain the following data:
- Tar Files: High-quality
.jpgfiles that passed aesthetic and NSFW filtering. - Parquet Files: Metadata for each tar file, including image paths, IDs, and processing scores.
- Naming Convention: Hash-based prefixes, such as
images-a1b2c3d4e5f6-000000.tar, keep distributed outputs unique. - Scores: The Parquet files store
aesthetic_scoreandnsfw_scorevalues.
Complete Tutorial
The command in Quickstart Step 3 downloads and runs the complete example, including its data and configuration options.
Next Steps
Continue with these image curation references and the full tutorial: