Qwen3-VL-4B-Instruct

View as Markdown

Qwen3-VL is Alibaba Cloud’s third-generation vision language model series. The MoE variant activates a fraction of parameters per token for efficient large-scale inference.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Qwen3-VL-4B-Instruct

From the repository root, run:

uv run automodel examples/vlm_finetune/qwen3/qwen3_vl_4b_instruct_rdr.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Qwen3-VL 4B on RDR ItemsUse qwen3_vl_4b_instruct_rdr.yaml. Dataset: rdr-items.

Fine-Tuning

See the VLM Fine-Tuning Guide.

Dense Qwen3-VL supports context parallelism. To distribute the vision tower across the CP group, enable distributed.multimodal.vision.frame_sharding; see the Context-Parallel Vision Frame Sharding guide. Qwen3-VL-MoE does not yet support this path.

Model Reference

Model Architecture

PropertyValue
TaskImage-Text-to-Text
ArchitectureQwen3VLForConditionalGeneration
Parameters4B
Hugging Face OrganizationQwen

Available Models

ModelHF ID
Qwen3-VL 4B InstructQwen/Qwen3-VL-4B-Instruct
Qwen3-VL-4B-ThinkingQwen/Qwen3-VL-4B-Thinking