> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# LLaVA-OneVision-1.5-4B-Instruct

> Use LLaVA-OneVision-1.5-4B-Instruct with NeMo AutoModel for vision-language fine-tuning, including checkpoints, runnable recipes, setup, and reference details.

[LLaVA-OneVision 1.5](https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2) is a vision-language model combining a **Rice ViT** encoder with a **Qwen3** language backbone, capable of handling both image and video understanding. NeMo AutoModel ships a custom NVIDIA implementation (`LlavaOneVisionForConditionalGeneration`) with FSDP2/HSDP support, LoRA fine-tuning and distributed training.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Fine-Tune LLaVA-OneVision-1.5-4B-Instruct

From the repository root, run:

```bash
uv run automodel --nproc-per-node=8 examples/vlm_finetune/llava_onevision/llava_ov_1_5_4b_finetune.yaml
```

## Choose a Workflow

| Goal                                                                         | Start Here                                                                                                                                                        |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Supervised fine-tuning (SFT) - LLaVA-OneVision-1.5 4B on LLaVA-Instruct-150K | Use [llava\_ov\_1\_5\_4b\_finetune.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/vlm_finetune/llava_onevision/llava_ov_1_5_4b_finetune.yaml). |

## Model Reference

### Model Architecture

| Property                  | Value                                       |
| ------------------------- | ------------------------------------------- |
| Task                      | Image-Text-to-Text                          |
| Architecture              | `LlavaOneVisionForConditionalGeneration`    |
| Parameters                | 4B                                          |
| Hugging Face Organization | [lmms-lab](https://huggingface.co/lmms-lab) |

* `LlavaOneVisionForConditionalGeneration`

Vision tower is the **Rice Transformer**: 14x14 patch embed with 2D RoPE, standard Transformer blocks (LayerNorm + Attention + MLP), and a 2x2 spatial Patch Merger that projects to the language-model hidden size.

### Available Models

| Model                           | HF ID                                                                                                         |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| LLaVA-OneVision-1.5 4B Instruct | [`lmms-lab/LLaVA-OneVision-1.5-4B-Instruct`](https://huggingface.co/lmms-lab/LLaVA-OneVision-1.5-4B-Instruct) |

## Related Resources

* [VLM Fine-Tuning Guide](/recipes-e2e-examples/gemma-3-3n)