Qwen-Image-2.1

View as Markdown

Qwen-Image-2.1 is Alibaba Cloud’s flow-matching text-to-image model. NeMo AutoModel supports full-parameter and LoRA fine-tuning of Qwen-Image-2.1.

Model Reference

Model Architecture

PropertyValue
TaskText-to-Image
ArchitectureDiffusion Transformer (Flow Matching)
HF OrgQwen

Available Models

Recipes

RecipeDescription
qwen_image_21_t2i_flow.yamlFull fine-tune of Qwen-Image-2.1 with FSDP2
qwen_image_21_t2i_flow_lora.yamlLoRA fine-tune of Qwen-Image-2.1
generate_qwen_image_21.yamlGenerate images with Qwen-Image-2.1

Overview

Qwen-Image-2.1 is a 7B single-stream, block-causal DiT conditioned on Qwen3-VL hidden states. It uses a 64-channel RGBA VAE with 16x spatial compression, so image sides must be multiples of 32 pixels. It requires a Diffusers build that provides QwenImage21Pipeline.

Prepare the Dataset

Cache VAE latents and Qwen3-VL prompt embeddings with the qwen_image_21 processor. RGB images receive an opaque alpha channel before VAE encoding:

uv run python -m tools.diffusion.preprocessing_multiprocess image \
--dataset_name lambdalabs/naruto-blip-captions \
--dataset_media_column image --dataset_caption_column text \
--processor qwen_image_21 \
--resolution_preset 1024p \
--verify \
--output_dir /cache/naruto-qwen-image-21

Fine-Tune Qwen-Image-2.1

Set data.dataloader.cache_dir and checkpoint.checkpoint_dir, then launch one FSDP2 rank per GPU:

uv run torchrun --nproc-per-node=8 \
examples/diffusion/finetune/finetune.py \
-c examples/diffusion/finetune/qwen_image_21_t2i_flow.yaml

The qwen_image_21 adapter supervises only the target-image tokens of the joint text and image sequence. The transformer lays out RoPE positions once per batch. Prompt padding would shift the image positions, so the adapter runs each sample as its own unpadded transformer call and sees the positions of single-prompt inference. Qwen-Image-2.1 is sampled without classifier-free guidance, so the recipes set cfg_dropout_prob: 0.0. On 8x H100 80 GB at 1024x1024 with a global batch of 16, the full fine-tune runs at about 1.4 s per step and uses about 60 GB per GPU; LoRA runs at about 1.05 s per step and uses about 56 GB per GPU.

Generate Images

Generate with the base model or a training checkpoint:

uv run python examples/diffusion/generate/generate.py \
-c examples/diffusion/generate/configs/generate_qwen_image_21.yaml \
--model.checkpoint /checkpoints/qwen-image-21/epoch_7_step_599