Quickstart for NeMo Curator Video Curation
Use this quickstart to install NeMo Curator, split videos into clips, and generate clip embeddings. Use the output for duplicate removal, similarity search, captioning, or model training.
Prerequisites
Prepare the software, GPU, and input data before you run a video pipeline:
- Ubuntu 20.04, 22.04, or 24.04; Python 3.11, 3.12, or 3.13;
uv; and Git. - An NVIDIA GPU from the Volta architecture or newer, with compute capability 7.0 or higher and CUDA 12 or later. CPU-only processing is not supported.
- VRAM requirements depend on the pipeline: approximately 16 GB for splitting and embedding, 38 GB for captioning with embedding, or 21 GB for reduced configurations with lower batch sizes and FP8.
- FFmpeg 8.0 or later with
h264_nvencorlibvpx-vp9. NVIDIA A100 and H100 GPUs do not include NVENC hardware. - Input videos in a local directory or an S3-compatible location.
Install uv if it is not already available. Refer to the NeMo Curator Installation Guide for setup details.
Some video operations, including target-resolution frame resizing and motion-vector
filtering, use OpenCV. The base installation does not include it; add
uv pip install "nemo-curator[cv2]" before using those operations. Motion-vector
decoding also requires the video_cpu or video_cuda12 extra, which provides
PyAV and its FFmpeg libraries. Stages that invoke the ffmpeg executable still
require the FFmpeg setup above; installing cv2 does not install that
executable or its codecs. If you use a container image, verify the active image
with python -c "import cv2; print(cv2.__version__)" and install the extra in
the container when needed.
Quickstart Steps
Complete these five steps to install Curator, verify FFmpeg, run the splitting example, and inspect its output.
-
Install NeMo Curator. A locally built container is recommended for video workflows so the maintained FFmpeg build and its encoders are available on every worker.
Build an Image (Recommended)
Source Installation
PyPI Installation
Follow Build an Image, then build the video image using Add FFmpeg for Audio or Video. Check out a published release tag first when you need a released version.
-
Verify that FFmpeg provides the encoder you plan to use and the software H.264 decoder required for CPU-side probing.
If an encoder or decoder is missing, follow Install FFmpeg and Encoders. Use
libvpx-vp9on GPUs without NVENC, including NVIDIA A100 and H100.
-
Create a model directory and set the input and output paths.
The container examples mount the host repository at
/workspace; keep input videos, output clips, and model files under that directory so they persist after the container exits. For source installs outside a container, replace these values with accessible host directories. Curator downloads selected model files toMODEL_DIRon first use. For S3 paths and credentials, refer to Data Directories and S3.
-
Run the video splitting example with a fixed 10-second stride.
Run this command from the Curator repository root. If you installed the package from PyPI, clone the Curator GitHub repository to access the example script. The script reads videos, splits them into clips, generates embeddings, and writes metadata.
Use
--transcode-encoder libvpx-vp9when your GPU does not include NVENC hardware.libopenh264requires a user-installed FFmpeg build. -
Confirm that the output directory contains clips, embeddings, and a manifest. Refer to Pipeline Output for the file layout and manifest example.
Reference Material
Use the following reference sections to configure storage, FFmpeg, models, and pipeline options.
Install FFmpeg and Encoders
Video pipelines use FFmpeg to decode and encode clips. Install FFmpeg 8.0 or later with an encoder supported by your GPU and a software H.264 decoder. The maintained script below builds FFmpeg with NVENC, libvpx-vp9, and software H.264/HEVC/AV1 decoders.
Debian/Ubuntu (Script)
Verify Installation
Run the maintained software decoder support script. It builds FFmpeg with NVIDIA NVENC, libvpx-vp9, and software H.264/HEVC/AV1 decoders in one step.
- Script source: software decoder support script.
Refer to Clip Encoding to select an encoder and verify NVENC support on your system.
On GPUs without NVENC hardware, such as NVIDIA A100 and NVIDIA H100, use --transcode-encoder libvpx-vp9. Software H.264 decoding is provided by install_h264_support.sh; the optional libopenh264 encoder is separate.
Available Models
Embeddings represent each video clip as a numeric vector that captures visual and semantic content. You can use these vectors to:
- Remove near-duplicate clips during duplicate removal
- Enable similarity search and clustering
- Support downstream analysis such as caption verification
Cosmos-Embed1 provides the model variants used for this quickstart:
Cosmos-Embed1
Cosmos-Embed1 is the default model. It has three variants: cosmos-embed1-224p, cosmos-embed1-336p, and cosmos-embed1-448p. The variants differ in input resolution and accuracy-to-VRAM tradeoffs. Curator downloads the selected model to MODEL_DIR on first use.
Download links for each model are available on Hugging Face:
- cosmos-embed1-224p on Hugging Face
- cosmos-embed1-336p on Hugging Face
- cosmos-embed1-448p on Hugging Face
This quickstart uses Cosmos-Embed1-224p.
Data Directories and S3
Use local paths for local processing. The quickstart uses DATA_DIR, OUT_DIR, and MODEL_DIR for input, output, and model locations.
For S3-compatible storage, configure credentials in ~/.aws/credentials and use s3:// paths for --video-dir and --output-clip-path. Apply these storage rules:
- Read input videos from S3 paths and write output clips to S3 paths.
- Keep model files local for performance.
- Confirm that your IAM permissions allow the required bucket reads and writes.
Configuration Options Reference
For complex configurations, store the command-line arguments in a file and pass the file path with the @ prefix:
The example uses fixed-stride splitting. The transnetv2 option supports scene-based splitting. The table lists the supported arguments:
Pipeline Output
After a successful run, the output directory contains:
The output directories contain these files:
- clips/ contains encoded video clips in MP4 format.
- embeddings/ contains NumPy arrays with clip embeddings for similarity search.
- metadata/manifest.jsonl contains clip paths, timestamps, and embedding paths.
- previews/ contains thumbnail images when you enable preview generation.
Example manifest entry:
Best Practices
Use these practices to prepare video data, select models, and monitor your pipeline:
Data Preparation
Validate and organize videos before processing:
- Validate input videos to detect corruption before processing.
- Convert videos to MP4 with H.264 for consistent results.
- Group similar videos to organize processing.
Model Selection
Choose a model variant that fits your accuracy and memory needs:
- Start with Cosmos-Embed1-224p for initial experiments.
- Use 336p or 448p when you need higher precision.
- Monitor GPU memory with
nvidia-smiduring processing.
Pipeline Configuration
Configure a representative run before processing a large dataset:
- Enable detailed logs with
--verbose. - Test the pipeline on five to 10 videos.
- Enable NVENC when your GPU supports it.
- Save embeddings and metadata for downstream tasks.
Infrastructure
Prepare storage and compute resources for your pipeline:
- Mount shared storage for multi-node processing.
- Plan for peak VRAM use during captioning and embedding.
- Monitor GPU use with
nvidia-smi dmon. - Schedule large video datasets as batch jobs.
Next Steps
Continue with the video curation guide and encoder reference: