Qwen-Image-2.1
Qwen-Image-2.1
Qwen-Image-2.1 is Alibaba Cloud’s flow-matching text-to-image model. NeMo AutoModel supports full-parameter and LoRA fine-tuning of Qwen-Image-2.1.
Model Reference
Model Architecture
Available Models
Recipes
Overview
Qwen-Image-2.1 is a 7B single-stream, block-causal DiT conditioned on Qwen3-VL hidden states. It uses a 64-channel RGBA VAE with 16x spatial compression, so image sides must be multiples of 32 pixels. It requires a Diffusers build that provides QwenImage21Pipeline.
Prepare the Dataset
Cache VAE latents and Qwen3-VL prompt embeddings with the qwen_image_21 processor. RGB images receive an opaque alpha channel before VAE encoding:
Fine-Tune Qwen-Image-2.1
Set data.dataloader.cache_dir and checkpoint.checkpoint_dir, then launch one FSDP2 rank per GPU:
The qwen_image_21 adapter supervises only the target-image tokens of the joint text and image sequence. The transformer lays out RoPE positions once per batch. Prompt padding would shift the image positions, so the adapter runs each sample as its own unpadded transformer call and sees the positions of single-prompt inference. Qwen-Image-2.1 is sampled without classifier-free guidance, so the recipes set cfg_dropout_prob: 0.0. On 8x H100 80 GB at 1024x1024 with a global batch of 16, the full fine-tune runs at about 1.4 s per step and uses about 60 GB per GPU; LoRA runs at about 1.05 s per step and uses about 56 GB per GPU.
Generate Images
Generate with the base model or a training checkpoint: