Model CoverageVision Language ModelsQwenQwen3-VL-235B-A22B-Instruct

Qwen3-VL-235B-A22B-Instruct

View as Markdown

Qwen3-VL-235B-A22B-Instruct has a checked-in NeMo AutoModel recipe for image-text-to-text. The Hugging Face configuration declares the Qwen3VLMoeForConditionalGeneration architecture.

Fine-Tune Qwen3-VL-235B-A22B-Instruct

Follow the installation instructions, then run the recipe from the repository root:

automodel examples/vlm_finetune/qwen3/qwen3_vl_moe_235b.yaml --nproc-per-node 8

Use the Slurm launcher guide for the multi-node run.

Choose a Workflow

GoalStart Here
Run the primary recipe for this modelUse the recipe configuration.
Try another checked-in recipe for this modelUse the alternate recipe.

Configuration

SettingConfiguration
Hardware-
StrategyFSDP2; tp_size=1, pp_size=4, cp_size=1, ep_size=32
Nodes32
FeaturesTransformer Engine, HybridEP (dispatcher=hybridep)
Advancedattention=sdpa, linear=te, experts=gmm, dispatcher=hybridep

More Recipes for This Model

Model Reference

Model Architecture

PropertyValue
TaskImage-text-to-text
ArchitectureQwen3VLMoeForConditionalGeneration

Available Models