Nemotron 3.5 Super VL#

Nemotron 3.5 Super VL combines a Nemotron-H hybrid language decoder with a RADIO vision encoder, a vision projector, and a separate temporal video embedder. It reuses the Nemotron Omni image/video stack without an audio encoder and retains shared Multi-Token Prediction (MTP) weights.

Megatron Bridge provides checkpoint conversion and pretraining, SFT, and PEFT recipes. The verification records below distinguish verified configurations from pending checks; adding a recipe does not imply that every workflow is verified.

The freeze_vision_model, freeze_vision_projection, and freeze_language_model provider options independently control which model components are trained. The Super-VL SFT recipes freeze the vision encoder while keeping the projection and language model trainable.

Verification scope: pretrain, SFT, PEFT, checkpoint resume, and dependent post-SFT export/inference remain unverified after the shared-provider callback fix. The commands, metrics, and artifact observations below are historical only: the copied provider configuration could omit training finalization callbacks, so finite losses and successful reloads do not establish synchronized, token-normalized updates. These runs also predate the current two-depth shared-MTP objective, FP32 optimizer-state and weight-decay defaults, and language-only PEFT targets. Pure conversion verification is unaffected. Historical throughput is not a corrected-training or optimized-performance baseline.

Verified configurations#

Choose an exact recorded configuration to see its command and expected result. These selectors are generated from the authoritative verification cards and never synthesize combinations.

Run a configuration#

Choose a workflow, precision, and exact recorded combination. The command and expected result update below.

Import · CPU

â—‹ Unverified
Hardware
not specified
Precision
BF16
Last verified
—
Exact command

No runnable command is recorded for this status.

Expected result

CPU-only import is not claimed in this card; the distributed GPU import below is the verified conversion path for this checkpoint.

Import · GPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-25
Exact command
Command
./scripts/conversion/convert.sh import --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model nvidia/NVIDIA-Nemotron-3.5-Super-120B-A12B --hf-revision e86197a3bad449de618a5835f26835ce770c6f10 --megatron-path work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron --torch-dtype bfloat16 --tp 1 --pp 1 --ep 8 --etp 1 --trust-remote-code --low-memory-save
Expected result

The pinned eight-GPU BF16 import exits successfully after mapping 6,281 parameters per rank and creates a TP1/PP1/EP8/ETP1 iter_0000000 checkpoint containing the complete language, MTP, image, projector, and temporal-video weights. The paired export reloads this saved checkpoint successfully.

Export · CPU

â—‹ Unverified
Hardware
not specified
Precision
BF16
Last verified
—
Exact command

No runnable command is recorded for this status.

Expected result

CPU-only export is not claimed in this card; the distributed GPU export below is the verified conversion path for this checkpoint.

Export · GPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-25
Exact command
Command
./scripts/conversion/convert.sh export --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model nvidia/NVIDIA-Nemotron-3.5-Super-120B-A12B --hf-revision e86197a3bad449de618a5835f26835ce770c6f10 --megatron-path work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron/iter_0000000 --hf-path work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-hf-export --torch-dtype bfloat16 --tp 1 --pp 1 --ep 8 --etp 1 --trust-remote-code --distributed-save --save-every-n-ranks 1
Expected result

Strict distributed export exits successfully in 63 indexed shards. All 43,078 source tensors match in keys, shapes, dtypes, and values, with maximum difference zero. AutoConfig, AutoProcessor, and AutoModelForImageTextToText reload the local-only export natively as NemotronH_Omni_Reasoning_V3 across all eight visible H100s.

Pretrain · H100

â—‹ Unverified
Hardware
H100
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 8 --gpus-per-node 8 --recipe nemotron_35_super_vl_pretrain_64gpu_h100_bf16_config --mode pretrain --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/vision-pretrained/iter_0000000 --max_steps 120 dataset.path=work/data/datacomp-525k-energon checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-pretrain --save_interval 60 logger.log_interval=1 logger.log_throughput=true optimizer.optimizer_cpu_offload=true optimizer.optimizer_offload_fraction=0.25 optimizer.overlap_cpu_optimizer_d2h_h2d=false model.recompute_granularity=full model.recompute_method=uniform model.recompute_num_layers=1 model.recompute_modules=null logger.log_device_memory_used=true logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-pretrain-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-29. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. Exactly 64 H100s completed all 120 optimizer steps from the learned vision-pretrained checkpoint on DataComp image-caption samples at TP1/PP2/CP1/EP32/ETP1 and GBS/MBS 1280/1. Natural dropless routing and full/uniform one-layer language recompute, selective vision recompute, and 25% optimizer-state CPU offload plus the recipe's BF16 precision-aware optimizer produced finite loss, zero skipped or NaN iterations, and the resolved ConfigContainer. Full optimizer/RNG checkpoints at steps 60 and 120 each contain all 64 distributed model shards and 32 Energon data-parallel states.

Pretrain · GB200

â—‹ Unverified
Hardware
GB200
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 16 --gpus-per-node 4 --recipe nemotron_35_super_vl_pretrain_64gpu_gb200_bf16_config --mode pretrain --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/vision-pretrained/iter_0000000 --max_steps 120 dataset.path=work/data/datacomp-525k-energon checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-pretrain --save_interval 60 logger.log_interval=1 logger.log_throughput=true logger.log_device_memory_used=true logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-pretrain-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-28. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. On exactly 64 GB200s within one NVL72 domain, complete all 120 optimizer steps from the learned vision-pretrained checkpoint on DataComp image-caption samples at TP2/PP1/CP1/EP64/ETP1 and GBS/MBS 512/1. Natural dropless routing, finite loss, zero skipped or NaN iterations, no active language or vision recompute, scoped Transformer Engine CUDA-graph capture, and the resolved ConfigContainer are required. Full optimizer/RNG checkpoints at steps 60 and 120 each contain all 64 distributed model shards and 32 Energon data-parallel states.

SFT · H100

â—‹ Unverified
Hardware
H100
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 8 --gpus-per-node 8 --recipe nemotron_35_super_vl_sft_64gpu_h100_bf16_config --mode sft --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron/iter_0000000 --max_steps 100 dataset.path=work/data/datacomp-525k-energon dataset.do_validation=false dataset.do_test=false validation.eval_iters=0 validation.eval_interval=0 scheduler.lr_decay_iters=100 checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-sft --save_interval 100 checkpoint.save_optim=false checkpoint.save_rng=false logger.log_interval=1 logger.log_throughput=true logger.log_device_memory_used=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-sft-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-28. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. Exactly 64 H100s completed 100 BF16 SFT optimizer steps from the immutable source checkpoint on 128,000 DataComp image-caption samples at TP1/PP2/CP1/EP32/ETP1 and GBS/MBS 1280/1. The shifted production loss masks contain 2,900,779 supervised tokens. Natural routing, HybridEP flex dispatch, fixed expert capacity, selective language recompute with vision recompute disabled, and the precision-aware optimizer produced finite loss with zero skipped or NaN iterations. The model-only step-100 checkpoint contains all 64 distributed shards and 32 Energon states and reloaded successfully. This dataset verifies image-caption batching; it does not contain video samples.

SFT · GB200

â—‹ Unverified
Hardware
GB200
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 16 --gpus-per-node 4 --recipe nemotron_35_super_vl_sft_64gpu_gb200_bf16_config --mode sft --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron/iter_0000000 --max_steps 100 dataset.path=work/data/datacomp-525k-energon dataset.do_validation=false dataset.do_test=false validation.eval_iters=0 validation.eval_interval=0 scheduler.lr_decay_iters=100 checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-sft --save_interval 100 checkpoint.save_optim=false checkpoint.save_rng=false logger.log_interval=1 logger.log_throughput=true logger.log_device_memory_used=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-sft-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-29. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. Exactly 64 GB200s within one NVL72 domain completed 100 BF16 SFT optimizer steps from the immutable source checkpoint on 128,000 DataComp image-caption samples at TP2/PP1/CP1/EP64/ETP1 and GBS/MBS 1280/1. The shifted production loss masks contain 2,900,779 supervised tokens. Natural routing, eager all-to-all dispatch, fixed expert capacity, and no active language or vision recompute produced finite loss with zero skipped or NaN iterations. The model-only step-100 checkpoint contains all 64 distributed shards and 32 Energon states and reloaded successfully. This dataset verifies image-caption batching; it does not contain video samples.

Long Context · H100

â—‹ Unverified
Hardware
H100
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command

No runnable command is recorded for this status.

Expected result

No H100 long-context SFT result is claimed in this card.

Long Context · GB200

â—‹ Unverified
Hardware
GB200
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command

No runnable command is recorded for this status.

Expected result

No GB200 long-context SFT result is claimed in this card.

LoRA · H100

â—‹ Unverified
Hardware
H100
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 2 --gpus-per-node 8 --recipe nemotron_35_super_vl_peft_16gpu_h100_bf16_config --mode lora --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron/iter_0000000 --max_steps 100 dataset.path=work/data/datacomp-525k-energon dataset.do_validation=false dataset.do_test=false validation.eval_iters=0 validation.eval_interval=0 checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-peft --save_interval 100 checkpoint.save_optim=false checkpoint.save_rng=false logger.log_interval=1 logger.log_throughput=true logger.log_device_memory_used=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/h100-peft-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-29. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. Exactly 16 H100s completed the rank-32, alpha-32, zero-dropout LoRA objective for 100 optimizer steps from the immutable source checkpoint on 1,600 DataComp image-caption samples at TP4/PP2/CP1/EP8/ETP1 and GBS/MBS 16/1. The shifted production loss masks contain 35,540 supervised tokens. Frozen base weights, natural routing, dropless all-to-all dispatch, and no active language or vision recompute produced finite loss with zero skipped or NaN iterations. The step-100 adapter contains all 16 distributed shards and two Energon states and reloaded successfully. This dataset does not contain video samples. This historical recipe used unqualified adapter targets that also matched vision and projector modules; it does not verify the current language-only adapter placement.

LoRA · GB200

â—‹ Unverified
Hardware
GB200
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --wait --nodes 4 --gpus-per-node 4 --recipe nemotron_35_super_vl_peft_16gpu_gb200_bf16_config --mode lora --pretrained_checkpoint work/model-verification/nemotron-3.5-super-vl-120b-a12b/gpu-megatron/iter_0000000 --max_steps 100 dataset.path=work/data/datacomp-525k-energon dataset.do_validation=false dataset.do_test=false validation.eval_iters=0 validation.eval_interval=0 checkpoint.load=null --save_dir work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-peft --save_interval 100 checkpoint.save_optim=false checkpoint.save_rng=false logger.log_interval=1 logger.log_throughput=true logger.log_device_memory_used=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/nemotron-3.5-super-vl-120b-a12b/gb200-peft-config/ConfigContainer.yaml
Expected result

Historical-only observations recorded on 2026-08-28. The command, metrics and artifact observations do not establish synchronized, token-normalized training updates. Recipe-specific reverification after the shared-provider callback fix is pending. Exactly 16 GB200s completed the rank-32, alpha-32, zero-dropout LoRA objective for 100 optimizer steps from the immutable source checkpoint on 1,600 DataComp image-caption samples at TP2/PP1/CP1/EP16/ETP1 and GBS/MBS 16/1. The shifted production loss masks contain 34,537 supervised tokens. Frozen base weights, natural routing, dropless all-to-all dispatch, and no active language or vision recompute produced finite loss with zero skipped or NaN iterations. The step-100 adapter contains all 16 distributed shards and eight Energon states and reloaded successfully. This dataset does not contain video samples. This historical recipe used unqualified adapter targets that also matched vision and projector modules; it does not verify the current language-only adapter placement.