Qwen3.5-4B
Qwen3.5-4B
Qwen3.5 is Alibaba Cloud’s unified vision-language model series, including dense and MoE variants for image and multimodal understanding tasks.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Qwen3.5-4B
From the repository root, run:
Choose a Workflow
Fine-Tuning
See the VLM Fine-Tuning Guide.
Validated Large-Model, Long-Context Scale
The Qwen3.5-MoE VLM training path has been validated at both 397-billion-parameter model scale and 128K context length. Qwen3.5-397B-A17B completed a 10-step end-to-end training run on 256 H100 GPUs with FSDP2, CP64, EP64, packed sequences, full activation checkpointing, and a trainable vision tower. The workload mixed text, image, and video data and included genuine examples of approximately 120K tokens.
The run maintained finite, decreasing loss, validating the training mechanics for this combined large-model and long-context regime. It is not a model convergence result. The public 397B recipe provides a starting configuration for that checkpoint; use the 122B EP8/CP32 recipe as the published 128K long-context reference.
Dense Qwen3.5 and Qwen3.5-MoE support context-parallel vision frame sharding. The Qwen3.5-MoE path composes expert and context parallelism with packed sequences; see the Context-Parallel Vision Frame Sharding guide.
Model Reference
Model Architecture
Qwen3_5ForConditionalGeneration- dense modelsQwen3_5MoeForConditionalGeneration- MoE models