Model CoverageVision Language ModelsLMMS LabLLaVA-OneVision-1.5-8B-Instruct

LLaVA-OneVision-1.5-8B-Instruct

View as Markdown

LLaVA-OneVision-1.5-8B-Instruct has a checked-in NeMo AutoModel recipe for image-text-to-text. The Hugging Face configuration declares the LLaVAOneVision1_5_ForConditionalGeneration architecture.

Fine-Tune LLaVA-OneVision-1.5-8B-Instruct

Follow the installation instructions, then run the recipe from the repository root:

automodel examples/vlm_finetune/llava_onevision/llava_ov_1_5_8b_lora.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Run the primary recipe for this modelUse the recipe configuration.
Prepare your environmentFollow the installation instructions.

Configuration

SettingConfiguration
Hardware-
StrategyFSDP2; tp_size=1, cp_size=1
Nodes-
FeaturesCheckpointing (enabled=true), LoRA
Advanced-

Model Reference

Model Architecture

PropertyValue
TaskImage-text-to-text
ArchitectureLLaVAOneVision1_5_ForConditionalGeneration

Available Models