Qwen3-VL-4B-Instruct
Qwen3-VL-4B-Instruct
Qwen3-VL is Alibaba Cloud’s third-generation vision language model series. The MoE variant activates a fraction of parameters per token for efficient large-scale inference.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Qwen3-VL-4B-Instruct
From the repository root, run:
Choose a Workflow
Fine-Tuning
See the VLM Fine-Tuning Guide.
Dense Qwen3-VL supports context parallelism. To distribute the vision tower across the CP
group, enable distributed.multimodal.vision.frame_sharding; see the
Context-Parallel Vision Frame Sharding guide. Qwen3-VL-MoE
does not yet support this path.