LLaVA-OneVision-1.5-8B-Instruct
LLaVA-OneVision-1.5-8B-Instruct
LLaVA-OneVision-1.5-8B-Instruct has a checked-in NeMo AutoModel recipe for image-text-to-text.
The Hugging Face configuration declares the LLaVAOneVision1_5_ForConditionalGeneration architecture.
Fine-Tune LLaVA-OneVision-1.5-8B-Instruct
Follow the installation instructions, then run the recipe from the repository root: