InternVL3_5-4B

View as Markdown

InternVL is a vision language model from Shanghai AI Laboratory (OpenGVLab), combining a large vision encoder with an InternLM language backbone for strong multimodal performance.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune InternVL3_5-4B

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/vlm_finetune/internvl/internvl_3_5_4b.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - InternVL3.5 4B on MedPixUse internvl_3_5_4b.yaml. Dataset: MedPix-VQA.

Model Reference

Model Architecture

PropertyValue
TaskImage-Text-to-Text
ArchitectureInternVLForConditionalGeneration
Parameters4B
Hugging Face OrganizationOpenGVLab

Available Models

ModelHF ID
InternVL3.5 4BOpenGVLab/InternVL3_5-4B
InternVL3.5 4B HFOpenGVLab/InternVL3_5-4B-hf