Qwen3.5-4B

View as Markdown

Qwen3.5 is Alibaba Cloud’s unified vision-language model series, including dense and MoE variants for image and multimodal understanding tasks.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Qwen3.5-4B

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/vlm_finetune/qwen3_5/qwen3_5_4b.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Qwen3.5-VL 4B on MedPixUse qwen3_5_4b.yaml. Dataset: MedPix-VQA.

Fine-Tuning

See the VLM Fine-Tuning Guide.

Validated Large-Model, Long-Context Scale

The Qwen3.5-MoE VLM training path has been validated at both 397-billion-parameter model scale and 128K context length. Qwen3.5-397B-A17B completed a 10-step end-to-end training run on 256 H100 GPUs with FSDP2, CP64, EP64, packed sequences, full activation checkpointing, and a trainable vision tower. The workload mixed text, image, and video data and included genuine examples of approximately 120K tokens.

The run maintained finite, decreasing loss, validating the training mechanics for this combined large-model and long-context regime. It is not a model convergence result. The public 397B recipe provides a starting configuration for that checkpoint; use the 122B EP8/CP32 recipe as the published 128K long-context reference.

Dense Qwen3.5 and Qwen3.5-MoE support context-parallel vision frame sharding. The Qwen3.5-MoE path composes expert and context parallelism with packed sequences; see the Context-Parallel Vision Frame Sharding guide.

Model Reference

Model Architecture

PropertyValue
TaskImage-Text-to-Text
ArchitectureQwen3_5ForConditionalGeneration, Qwen3_5MoeForConditionalGeneration
Parameters4B
Hugging Face OrganizationQwen
  • Qwen3_5ForConditionalGeneration - dense models
  • Qwen3_5MoeForConditionalGeneration - MoE models

Available Models

ModelHF ID
Qwen3.5 4BQwen/Qwen3.5-4B