HunyuanImage-3.0

View as Markdown

HunyuanImage-3.0 is Tencent’s text-to-image model. Unlike the diffusers-based models on this site, it is a single multimodal MoE decoder: the prompt tokens and the image tokens share one sequence, and the same transformer predicts the flow-matching velocity of the image latents. NeMo AutoModel ships its own implementation of the transformer, so the diffusion recipe trains it through the custom-model path with FSDP2 and expert parallelism.

Model Reference

Model Architecture

PropertyValue
TasksText-to-Image
ArchitectureMultimodal MoE decoder with flow-matching image head (80B total / 13B active)
Layers32, each with GQA attention (32 query / 8 KV heads, head dim 128) and 64 routed experts (top-8) plus one shared expert
Positions2D RoPE over the joint text / image sequence
Latents32 channels, 16x spatial compression (release VAE)
HF Orgtencent

Model Class

HunyuanImage3ForCausalMM (nemo_automodel/components/models/hunyuan_image3/) is registered for the checkpoint’s HunyuanImage3ForCausalMM architecture and hunyuan_image_3_moe model type, so the config loads without trust_remote_code. The decoder uses the shared MoE layer, which the MoE parallelizer shards with FSDP2 and expert parallelism; the released checkpoint is converted by HunyuanImage3StateDictAdapter (fused [up; gate] expert projections to grouped experts, wte / ln_f renames). The release’s VAE and vision encoder are not part of the training model.

Supported: FSDP2, expert parallelism, LoRA, full fine-tuning through the same recipe. Not supported: tensor, pipeline and context parallelism, image conditioning (image-to-image), text generation training.

Available Models

ModelHF IDTraining Scope
HunyuanImage-3.0tencent/HunyuanImage-3.0Text-to-image LoRA and full fine-tuning

Recipes

RecipeDescription
hunyuan_image3_t2i_flow_lora.yamlSingle-node LoRA fine-tuning with 8-way expert parallelism
hunyuan_image3_t2i_flow.yamlFour-node full fine-tuning with 8-way expert parallelism
generate_hunyuan_image3.yamlText-to-image generation from the base model, a full fine-tune or a LoRA adapter

Prepare Data

The preprocessing step needs the release’s VAE and tokenizer wrapper, which are loaded from the checkpoint’s remote code (nemo_automodel/components/models/hunyuan_image3/release.py, shared with generation). For every image it stores the VAE latent and the release’s prompt token ids around the image span (the conditional prompt, the <cfg> unconditional prompt used for classifier-free guidance dropout, and the suffix). Images are snapped to the release’s own resolution group (33 aspect ratios around 1024x1024) instead of the generic buckets:

uv run python -m tools.diffusion.preprocessing_multiprocess image \
--image_dir /data/images \
--output_dir /cache/hunyuan-image3 \
--processor hunyuan_image3 \
--model_name tencent/HunyuanImage-3.0 \
--resolution_preset 1024p

Captions follow the image processors’ JSONL convention (<stem>_internvl.json next to the image, or the --caption_field you pass).

Cache Contract

Each cache sample contains:

  • latent: [32, height / 16, width / 16] (fp32, scaled by the VAE scaling factor)
  • prompt_input_ids: token ids before the image span, ending in <timestep>
  • uncond_prompt_input_ids: the same with the prompt replaced by <cfg> tokens
  • prompt_suffix_ids: token ids after the image span (<eoi>)
  • bucket_resolution, original_resolution, crop_offset, prompt, image_path

The hunyuan_image3 flow-matching adapter assembles prefix + <img> * (h * w) + suffix per sample and right-pads the batch; flow_matching.cfg_dropout_prob swaps in the unconditional prefix.

Train

Set data.dataloader.cache_dir and checkpoint.checkpoint_dir, then launch one rank per GPU. LoRA fits on one node:

uv run torchrun --nproc-per-node=8 \
examples/diffusion/finetune/finetune.py \
-c examples/diffusion/finetune/hunyuan_image3_t2i_flow_lora.yaml

Full fine-tuning keeps fp32-precision master weights and Adam moments for all 80B parameters (about 1.1 TB), so its recipe is laid out for four nodes; run the same command on every node with its own --node-rank:

uv run torchrun --nnodes=4 --node-rank=NODE_RANK --nproc-per-node=8 \
--master-addr=MASTER_ADDR --master-port=29500 \
examples/diffusion/finetune/finetune.py \
-c examples/diffusion/finetune/hunyuan_image3_t2i_flow.yaml

Both recipes keep the router in fp32 (gate_precision: float32), train with flow_shift: 3.0 (the value the release samples with) and use Transformer Engine’s FusedAdam with fp32 master weights (master_weights: true with 16-bit remainders), since plain AdamW on bf16 parameters rounds away most small updates. The LoRA recipe adapts the attention projections and the shared expert, and its checkpoints hold only the adapter (adapter_model.safetensors plus its config). fsdp.ep_size: 1 is also supported; the experts are then sharded by FSDP2.

Generate

examples/diffusion/generate/generate.py samples with the native transformer, sharded across the launched ranks as in training, and the release’s VAE, prompt format and flow_shift from the checkpoint. Sampling follows the release pipeline: Euler steps on the shifted sigma schedule with classifier-free guidance over a [prompt, <cfg>] batch. Every rank samples the same prompts:

uv run torchrun --nproc-per-node=8 \
examples/diffusion/generate/generate.py \
-c examples/diffusion/generate/configs/generate_hunyuan_image3.yaml

Set model.lora_weights to the model/ directory of a LoRA checkpoint step, or model.checkpoint to a full fine-tuning checkpoint step; both load through the training checkpointer. Requested sizes snap to the release resolution group.

Install

From a source checkout, install the diffusion media dependencies:

uv sync --locked --all-groups --extra all --extra diffusion-media

See the Installation Guide and Diffusion Training and Fine-Tuning Guide.