Model CoverageOmniNVIDIANVIDIA-Nemotron-3.5-Super-VL-120B-A12B-BF16

NVIDIA-Nemotron-3.5-Super-VL-120B-A12B-BF16

View as Markdown

NVIDIA-Nemotron-3.5-Super-VL-120B-A12B-BF16 pairs a Nemotron Super NemotronH (hybrid Mamba-2 and Attention) Mixture-of-Experts (MoE) language backbone with a RADIO v2.5-H vision encoder, sharing its NemotronH_Omni_Reasoning_V3 architecture and vision stack with Nemotron-3-Nano-Omni. The BF16 checkpoint has no audio encoder, so audio inputs are not supported.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune NVIDIA-Nemotron-3.5-Super-VL-120B-A12B-BF16

From the repository root, run:

uv run automodel examples/vlm_finetune/nemotron_3_5_super_vl/nemotron_3_5_super_vl_cord_v2_peft.yaml --nproc-per-node 8

The Tulu-3 text recipe runs on 4 nodes:

uv run automodel examples/vlm_finetune/nemotron_3_5_super_vl/nemotron_3_5_super_vl_tulu3_text_100steps.yaml --nnodes 4 --nproc-per-node 8

Choose a Workflow

GoalStart Here
Full SFT, text-only (vision/audio towers frozen and unused)Use nemotron_3_5_super_vl_tulu3_text_100steps.yaml. Dataset: Tulu-3.
Context-parallel parity test, cp_size=8; text-only packs (inject_fake_images: false)Use nemotron_3_5_super_vl_tulu3_packed_cp8_100steps.yaml. Dataset: Tulu-3 (packed, THD).
Low-rank adaptation (LoRA) (rank 64) on the LLM linear projections, base model frozen; receipt image to structured field tokens; single nodeUse nemotron_3_5_super_vl_cord_v2_peft.yaml. Dataset: CORD-v2.

LoRA on CORD-v2

nemotron_3_5_super_vl_cord_v2_peft.yaml fine-tunes the model on CORD-v2 receipt field extraction (the same task as the NemotronOmni CORD-v2 guide) with LoRA adapters instead of full SFT. Because the base weights are frozen, there are no fp32 master weights or optimizer states for them, so the 121B model trains on a single node of 8 H100 (cp_size 1, ep_size 8) instead of the 4 nodes that full SFT needs.

PropertyValue
Low-rank adaptation (LoRA) targets272 linear projections: Mamba in_proj/out_proj, attention q/k/v/o_proj, MoE fc1/fc2_latent_proj and shared experts (routed experts, vision tower, projector, lm_head frozen)
Trainable parameters177M (0.15%), rank 64, alpha 128; TE FusedAdam with fp32 master weights, lr 2e-4
Peak memory38.8 GiB/GPU
Adapter checkpoint355 MB (adapter_model.safetensors + adapter_config.json)
Training400 steps (global batch 8, 4 epochs of the 800 training receipts) in 13 min on one node; train loss 0.51 -> 0.01
Validation loss0.041 / 0.039 / 0.043 / 0.040 at steps 99 / 199 / 299 / 399
Field extraction (first 20 validation receipts, greedy, adapter merged into the base model)step 199 (LOWEST_VAL): 8/20 exact match, mean similarity 0.954 (median 0.997), 20/20 structured outputs

Model Reference

Model Architecture

PropertyValue
TaskMultimodal (Text/Image)
ArchitectureNemotronH_Omni_Reasoning_V3
Parameters121B total (MoE, 512 routed experts, 22 active per token; 124B with the MTP head)
Hugging Face Organizationnvidia
  • NemotronH_Omni_Reasoning_V3 reuses the same NemotronOmniForConditionalGeneration wrapper, vision tower, and state-dict adapter as NemotronH_Nano_Omni_Reasoning_V3; only the LLM sub-config scale differs (Super compared to Nano). The sound_encoder and sound_projection submodules are constructed only when the checkpoint’s sound_config is present, so this checkpoint loads with audio disabled.

Available Models