LLaVA-OneVision-1.5-4B-Instruct
LLaVA-OneVision-1.5-4B-Instruct
LLaVA-OneVision 1.5 is a vision-language model combining a Rice ViT encoder with a Qwen3 language backbone, capable of handling both image and video understanding. NeMo AutoModel ships a custom NVIDIA implementation (LlavaOneVisionForConditionalGeneration) with FSDP2/HSDP support, LoRA fine-tuning and distributed training.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune LLaVA-OneVision-1.5-4B-Instruct
From the repository root, run:
Choose a Workflow
Model Reference
Model Architecture
LlavaOneVisionForConditionalGeneration
Vision tower is the Rice Transformer: 14x14 patch embed with 2D RoPE, standard Transformer blocks (LayerNorm + Attention + MLP), and a 2x2 spatial Patch Merger that projects to the language-model hidden size.