Get Started

Quickstart for NeMo Curator Video Curation

View as Markdown

Use this quickstart to install NeMo Curator, split videos into clips, and generate clip embeddings. Use the output for duplicate removal, similarity search, captioning, or model training.

Prerequisites

Prepare the software, GPU, and input data before you run a video pipeline:

  • Ubuntu 20.04, 22.04, or 24.04; Python 3.11, 3.12, or 3.13; uv; and Git.
  • An NVIDIA GPU from the Volta architecture or newer, with compute capability 7.0 or higher and CUDA 12 or later. CPU-only processing is not supported.
  • VRAM requirements depend on the pipeline: approximately 16 GB for splitting and embedding, 38 GB for captioning with embedding, or 21 GB for reduced configurations with lower batch sizes and FP8.
  • FFmpeg 8.0 or later with h264_nvenc or libvpx-vp9. NVIDIA A100 and H100 GPUs do not include NVENC hardware.
  • Input videos in a local directory or an S3-compatible location.

Install uv if it is not already available. Refer to the NeMo Curator Installation Guide for setup details.

Some video operations, including target-resolution frame resizing and motion-vector filtering, use OpenCV. The base installation does not include it; add uv pip install "nemo-curator[cv2]" before using those operations. Motion-vector decoding also requires the video_cpu or video_cuda12 extra, which provides PyAV and its FFmpeg libraries. Stages that invoke the ffmpeg executable still require the FFmpeg setup above; installing cv2 does not install that executable or its codecs. If you use a container image, verify the active image with python -c "import cv2; print(cv2.__version__)" and install the extra in the container when needed.

curl -LsSf https://astral.sh/uv/0.12.1/install.sh | sh
source $HOME/.local/bin/env

Quickstart Steps

Complete these five steps to install Curator, verify FFmpeg, run the splitting example, and inspect its output.

  1. Install NeMo Curator. A locally built container is recommended for video workflows so the maintained FFmpeg build and its encoders are available on every worker.

  2. Verify that FFmpeg provides the encoder you plan to use and the software H.264 decoder required for CPU-side probing.

    ffmpeg -hide_banner -version | head -n 5
    ffmpeg -encoders | grep -E "h264_nvenc|libvpx-vp9"
    ffmpeg -decoders | grep -E '^[[:space:]]+V[^[:space:]]+[[:space:]]+h264[[:space:]]'

    If an encoder or decoder is missing, follow Install FFmpeg and Encoders. Use libvpx-vp9 on GPUs without NVENC, including NVIDIA A100 and H100.

  1. Create a model directory and set the input and output paths.

    DATA_DIR=/workspace/videos
    OUT_DIR=/workspace/output_clips
    MODEL_DIR=/workspace/models
    mkdir -p "$MODEL_DIR"

    The container examples mount the host repository at /workspace; keep input videos, output clips, and model files under that directory so they persist after the container exits. For source installs outside a container, replace these values with accessible host directories. Curator downloads selected model files to MODEL_DIR on first use. For S3 paths and credentials, refer to Data Directories and S3.

  1. Run the video splitting example with a fixed 10-second stride.

    Run this command from the Curator repository root. If you installed the package from PyPI, clone the Curator GitHub repository to access the example script. The script reads videos, splits them into clips, generates embeddings, and writes metadata.

    python tutorials/video/getting-started/video_split_clip_example.py \
    --video-dir "$DATA_DIR" \
    --model-dir "$MODEL_DIR" \
    --output-clip-path "$OUT_DIR" \
    --splitting-algorithm fixed_stride \
    --fixed-stride-split-duration 10.0 \
    --embedding-algorithm cosmos-embed1-224p \
    --transcode-encoder h264_nvenc \
    --verbose

    Use --transcode-encoder libvpx-vp9 when your GPU does not include NVENC hardware. libopenh264 requires a user-installed FFmpeg build.

  2. Confirm that the output directory contains clips, embeddings, and a manifest. Refer to Pipeline Output for the file layout and manifest example.

Reference Material

Use the following reference sections to configure storage, FFmpeg, models, and pipeline options.

Install FFmpeg and Encoders

Video pipelines use FFmpeg to decode and encode clips. Install FFmpeg 8.0 or later with an encoder supported by your GPU and a software H.264 decoder. The maintained script below builds FFmpeg with NVENC, libvpx-vp9, and software H.264/HEVC/AV1 decoders.

Run the maintained software decoder support script. It builds FFmpeg with NVIDIA NVENC, libvpx-vp9, and software H.264/HEVC/AV1 decoders in one step.

curl -fsSL https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/docker/common/install_h264_support.sh -o install_h264_support.sh
chmod +x install_h264_support.sh
sudo bash install_h264_support.sh

Refer to Clip Encoding to select an encoder and verify NVENC support on your system.

On GPUs without NVENC hardware, such as NVIDIA A100 and NVIDIA H100, use --transcode-encoder libvpx-vp9. Software H.264 decoding is provided by install_h264_support.sh; the optional libopenh264 encoder is separate.

Available Models

Embeddings represent each video clip as a numeric vector that captures visual and semantic content. You can use these vectors to:

  • Remove near-duplicate clips during duplicate removal
  • Enable similarity search and clustering
  • Support downstream analysis such as caption verification

Cosmos-Embed1 provides the model variants used for this quickstart:

Cosmos-Embed1

Cosmos-Embed1 is the default model. It has three variants: cosmos-embed1-224p, cosmos-embed1-336p, and cosmos-embed1-448p. The variants differ in input resolution and accuracy-to-VRAM tradeoffs. Curator downloads the selected model to MODEL_DIR on first use.

Model VariantResolutionVRAM UsageSpeedAccuracyBest For
cosmos-embed1-224p224 × 224~8 GBFastestGoodLarge-scale processing, initial curation
cosmos-embed1-336p336 × 336~12 GBMediumBetterBalanced performance and quality
cosmos-embed1-448p448 × 448~16 GBSlowerBestHigh-quality embeddings, fine-grained matching

Download links for each model are available on Hugging Face:

This quickstart uses Cosmos-Embed1-224p.

Data Directories and S3

Use local paths for local processing. The quickstart uses DATA_DIR, OUT_DIR, and MODEL_DIR for input, output, and model locations.

For S3-compatible storage, configure credentials in ~/.aws/credentials and use s3:// paths for --video-dir and --output-clip-path. Apply these storage rules:

  • Read input videos from S3 paths and write output clips to S3 paths.
  • Keep model files local for performance.
  • Confirm that your IAM permissions allow the required bucket reads and writes.

Configuration Options Reference

For complex configurations, store the command-line arguments in a file and pass the file path with the @ prefix:

echo '--video-dir /data/videos
--output-clip-path /data/output
--splitting-algorithm fixed_stride
--fixed-stride-split-duration 10.0
--embedding-algorithm cosmos-embed1-224p
--transcode-encoder h264_nvenc' > my_config.txt
python tutorials/video/getting-started/video_split_clip_example.py @my_config.txt

The example uses fixed-stride splitting. The transnetv2 option supports scene-based splitting. The table lists the supported arguments:

OptionValuesDescription
Splitting  
--splitting-algorithmfixed_stride, transnetv2Method for dividing videos into clips
--fixed-stride-split-durationFloat (seconds)Clip length for fixed stride (default: 10.0)
--transnetv2-frame-decoder-modepynvc, ffmpeg_gpu, ffmpeg_cpuFrame decoding method for TransNetV2
Embedding  
--embedding-algorithmcosmos-embed1-224p, cosmos-embed1-336p, cosmos-embed1-448pEmbedding model to use
Encoding  
--transcode-encoderh264_nvenc, libvpx-vp9, libopenh264Video encoder for output clips. Use libvpx-vp9 (CPU) on GPUs without NVENC such as A100/H100. libopenh264 is accepted but requires a user-installed FFmpeg build. Refer to software codec support.
--transcode-use-hwaccelFlagEnable hardware acceleration for encoding (only valid with h264_nvenc).
Optional Features  
--generate-captionsFlagGenerate text captions for each clip
--generate-previewsFlagCreate preview images for each clip
--verboseFlagEnable detailed logging output

Pipeline Output

After a successful run, the output directory contains:

$OUT_DIR/
├── clips/
│ ├── video1_clip_0000.mp4
│ ├── video1_clip_0001.mp4
│ └── ...
├── embeddings/
│ ├── video1_clip_0000.npy
│ ├── video1_clip_0001.npy
│ └── ...
├── metadata/
│ └── manifest.jsonl
└── previews/ (if --generate-previews enabled)
├── video1_clip_0000.jpg
└── ...

The output directories contain these files:

  • clips/ contains encoded video clips in MP4 format.
  • embeddings/ contains NumPy arrays with clip embeddings for similarity search.
  • metadata/manifest.jsonl contains clip paths, timestamps, and embedding paths.
  • previews/ contains thumbnail images when you enable preview generation.

Example manifest entry:

{
"video_path": "/data/input_videos/video1.mp4",
"clip_path": "/data/output_clips/clips/video1_clip_0000.mp4",
"start_time": 0.0,
"end_time": 10.0,
"embedding_path": "/data/output_clips/embeddings/video1_clip_0000.npy",
"preview_path": "/data/output_clips/previews/video1_clip_0000.jpg"
}

Best Practices

Use these practices to prepare video data, select models, and monitor your pipeline:

Data Preparation

Validate and organize videos before processing:

  • Validate input videos to detect corruption before processing.
  • Convert videos to MP4 with H.264 for consistent results.
  • Group similar videos to organize processing.

Model Selection

Choose a model variant that fits your accuracy and memory needs:

  • Start with Cosmos-Embed1-224p for initial experiments.
  • Use 336p or 448p when you need higher precision.
  • Monitor GPU memory with nvidia-smi during processing.

Pipeline Configuration

Configure a representative run before processing a large dataset:

  • Enable detailed logs with --verbose.
  • Test the pipeline on five to 10 videos.
  • Enable NVENC when your GPU supports it.
  • Save embeddings and metadata for downstream tasks.

Infrastructure

Prepare storage and compute resources for your pipeline:

  • Mount shared storage for multi-node processing.
  • Plan for peak VRAM use during captioning and embedding.
  • Monitor GPU use with nvidia-smi dmon.
  • Schedule large video datasets as batch jobs.

Next Steps

Continue with the video curation guide and encoder reference: