Evaluate Video Caption Quality
Use the repository’s Summarize-then-Align tools to build a fixed, diverse video benchmark and detect large semantic regressions in generated captions.
This workflow is evaluation tooling, not a stage in the main NeMo Curator pipeline. It reads artifacts produced by video pipelines and writes benchmark symlinks, cached summaries, and a score CSV. It does not filter or modify a curated dataset.
How Summarize-then-Align Works
Long captions can exceed the CosmosEmbed1 text encoder’s context and can receive different scores merely because one model is more verbose. The scorer therefore:
- Reads and concatenates the window captions for each clip.
- Uses a summarizer LLM to retain visible objects, actions, positions, colors, clothing, and on-screen text in fewer than 80 words.
- Encodes the summary with the text side of CosmosEmbed1.
- Computes cosine similarity between the text embedding and the precomputed video embedding.
- Reports the mean for each captioning model and writes every per-clip score to CSV.
Use the score as a fast regression signal, not as an absolute quality judgment. It measures semantic alignment in one embedding space; it does not directly measure fluency, completeness, temporal precision, hallucination severity, or usefulness for a downstream task.
Before You Start
- Run from a source checkout because
eval/video/build_benchmark_dataset.pyandeval/video/caption_clipscore.pyare repository tools. - Use a Linux host with an NVIDIA CUDA GPU. Benchmark construction runs the GPU video pipeline, and scoring loads the summarizer and CosmosEmbed1 on CUDA in BF16.
- Install the
video_cuda12environment described in Get Started with Video Curation. The environment includes vLLM, Transformers, PyTorch, and the video dependencies; the development groups provide scikit-learn for K-means. - Provide enough local storage for the sampled pipeline output. The final benchmark uses symlinks, but
_pipeline_outputretains generated clips, metadata, and embeddings. - Obtain the model assets before an offline run. The scorer uses local files only for CosmosEmbed1.
The supported CosmosEmbed1 variants are:
The model root must contain the Hugging Face organization and repository directories. For example, --cosmos-model-dir /models with --variant 224p resolves /models/nvidia/Cosmos-Embed1-224p.
build_benchmark_dataset.py always generates 224p embeddings. Keep the scorer’s default --variant 224p for those artifacts. Supplying a different variant does not convert existing embeddings and can cause incompatible vector shapes.
Step 1: Build a Diverse Benchmark
The builder samples source videos, creates fixed-stride clips, filters them by aesthetic score, computes 224p embeddings, clusters all surviving clips with K-means, and selects the clip nearest each centroid. It tries to select only one clip per source video; when a cluster has no unused source, it falls back to its nearest clip.
The input can be OpenVid-1M or another dataset you are licensed to use. The script scans only the top level of --video-dir for filenames ending in lowercase .mp4; discovery is not recursive.
From the repository root, run:
Builder Arguments
Benchmark Layout
Each line of selected_uids.txt is tab-separated:
The source basename and duration span let the scorer resolve a selected clip after a new pipeline run generates different clip UUIDs.
The benchmark directories contain absolute symlinks and are not portable by themselves. Preserve _pipeline_output and the original source videos, or materialize the links before moving the benchmark to another machine. For a clean rebuild, use a new --output-dir; existing links are skipped, and stale pipeline artifacts are not removed automatically.
Step 2: Generate Comparable Captions
Run every captioning model with the same benchmark inputs, split duration, filtering, and embedding variant. The current video example generates embeddings by default, so no separate --generate-embeddings flag is needed.
Repeat the command with a new output path and another supported --captioning-algorithm, such as qwen3, nemotron-bf16, nemotron-fp8, nemotron-nvfp4, or nemotron-3-nano-omni. Refer to Captions and Preview for caption-model configuration.
The scorer expects each caption output directory to contain metas/v0/<uuid>.json. A minimal record looks like:
For each window, the scorer uses the first nonblank string whose key contains caption. The LABEL in --caption-dirs LABEL=PATH names a result column; it does not select a metadata key. Use an output directory containing only the caption field you intend to evaluate, rather than mixing original and enhanced caption fields in one record.
Step 3: Score and Cache Summaries
Choose one caption run’s ce1_embd/ directory as the shared video-embedding source. The scorer evaluates only UUIDs present in that directory and in every caption directory, then applies the optional selected-UID filter.
Use a summarizer from a different model family than the captioners to reduce vocabulary and phrasing bias:
The summarizer runs first with temperature 0, a 120-token generation limit, BF16 weights, and 50% vLLM GPU-memory utilization. It is unloaded before CosmosEmbed1 is loaded, so both models do not need to reside on the GPU simultaneously.
Scorer Arguments
To rerun the same comparison without loading the summarizer:
The summary cache is nested by clip UUID and label:
Cache labels and caption contents together as one evaluation fixture. If you rename a label, change captions, or add a model, regenerate the affected summaries. A missing cache entry becomes an empty string and produces a warning, but scoring continues; treat that result as invalid.
Outputs and Interpretation
The CSV contains the resolved clip UUID, the source video reported by the first caption directory, and one score column per label:
Scores are cosine similarities with a theoretical range of -1 to 1. Compare means only across runs that use the same benchmark, video embeddings, CosmosEmbed1 variant, summarizer prompt/model, cache policy, and caption-window extraction behavior.
For the original 200-clip OpenVid-1M benchmark, observed means were approximately 0.39 for Qwen2.5-VL and Qwen3-VL and 0.38 for Nemotron-VL. These values are reference observations, not universal thresholds. A practical regression workflow is:
- Keep the benchmark and summaries fixed.
- Flag a mean decrease greater than 5% relative to your own accepted baseline.
- Sort the new model’s CSV column and review at least the ten lowest-scoring clips.
- Decide from human review whether the change is a real semantic regression or an artifact of summary style, verbosity, or the embedding model.
Reproducibility and Failure Modes
--seedcontrols Python sampling and the K-means random state, but library versions, model revisions, decoding, GPU kernels, and source-file ordering can still affect the selected benchmark. Version the resultingselected_uids.txtrather than rebuilding it for every comparison.CUBLAS_WORKSPACE_CONFIG=:4096:8, temperature 0, and cached summaries reduce variation, but do not guarantee bitwise-identical scores across different hardware or software stacks.- The scorer uses the intersection of embeddings and all caption metadata directories. Missing outputs silently reduce the evaluated set; compare the logged
Evaluating N clipsvalue with the expected benchmark size. - Source/span fallback reads metadata only from the first
--caption-dirspath. Pass the most complete caption run first. A benchmark clip missing there cannot be resolved from a later caption directory. - Source/span fallback uses the source-video basename and timestamps rounded to four decimal places. Duplicate basenames or changed split settings can resolve incorrectly or not at all.
- Empty caption windows are concatenated to an empty caption and are still summarized. Inspect metadata before treating a low score as a model-quality result.
--num-clustersmust not exceed the number of embeddings left after splitting and aesthetic filtering. Lower the cluster count, sample more videos, or relax filtering if K-means rejects the input.- Reuse a single embedding directory only when every caption run represents the same clips and CosmosEmbed1 variant.