HunyuanImage-3.0
HunyuanImage-3.0
HunyuanImage-3.0 is Tencent’s text-to-image model. Unlike the diffusers-based models on this site, it is a single multimodal MoE decoder: the prompt tokens and the image tokens share one sequence, and the same transformer predicts the flow-matching velocity of the image latents. NeMo AutoModel ships its own implementation of the transformer, so the diffusion recipe trains it through the custom-model path with FSDP2 and expert parallelism.
Model Reference
Model Architecture
Model Class
HunyuanImage3ForCausalMM (nemo_automodel/components/models/hunyuan_image3/) is registered for the checkpoint’s HunyuanImage3ForCausalMM architecture and hunyuan_image_3_moe model type, so the config loads without trust_remote_code. The decoder uses the shared MoE layer, which the MoE parallelizer shards with FSDP2 and expert parallelism; the released checkpoint is converted by HunyuanImage3StateDictAdapter (fused [up; gate] expert projections to grouped experts, wte / ln_f renames). The release’s VAE and vision encoder are not part of the training model.
Supported: FSDP2, expert parallelism, LoRA, full fine-tuning through the same recipe. Not supported: tensor, pipeline and context parallelism, image conditioning (image-to-image), text generation training.
Available Models
Recipes
Prepare Data
The preprocessing step needs the release’s VAE and tokenizer wrapper, which are loaded from the checkpoint’s remote code (nemo_automodel/components/models/hunyuan_image3/release.py, shared with generation). For every image it stores the VAE latent and the release’s prompt token ids around the image span (the conditional prompt, the <cfg> unconditional prompt used for classifier-free guidance dropout, and the suffix). Images are snapped to the release’s own resolution group (33 aspect ratios around 1024x1024) instead of the generic buckets:
Captions follow the image processors’ JSONL convention (<stem>_internvl.json next to the image, or the --caption_field you pass).
Cache Contract
Each cache sample contains:
latent:[32, height / 16, width / 16](fp32, scaled by the VAE scaling factor)prompt_input_ids: token ids before the image span, ending in<timestep>uncond_prompt_input_ids: the same with the prompt replaced by<cfg>tokensprompt_suffix_ids: token ids after the image span (<eoi>)bucket_resolution,original_resolution,crop_offset,prompt,image_path
The hunyuan_image3 flow-matching adapter assembles prefix + <img> * (h * w) + suffix per sample and right-pads the batch; flow_matching.cfg_dropout_prob swaps in the unconditional prefix.
Train
Set data.dataloader.cache_dir and checkpoint.checkpoint_dir, then launch one rank per GPU. LoRA fits on one node:
Full fine-tuning keeps fp32-precision master weights and Adam moments for all 80B parameters (about 1.1 TB), so its recipe is laid out for four nodes; run the same command on every node with its own --node-rank:
Both recipes keep the router in fp32 (gate_precision: float32), train with flow_shift: 3.0 (the value the release samples with) and use Transformer Engine’s FusedAdam with fp32 master weights (master_weights: true with 16-bit remainders), since plain AdamW on bf16 parameters rounds away most small updates. The LoRA recipe adapts the attention projections and the shared expert, and its checkpoints hold only the adapter (adapter_model.safetensors plus its config). fsdp.ep_size: 1 is also supported; the experts are then sharded by FSDP2.
Generate
examples/diffusion/generate/generate.py samples with the native transformer, sharded across the launched ranks as in training, and the release’s VAE, prompt format and flow_shift from the checkpoint. Sampling follows the release pipeline: Euler steps on the shifted sigma schedule with classifier-free guidance over a [prompt, <cfg>] batch. Every rank samples the same prompts:
Set model.lora_weights to the model/ directory of a LoRA checkpoint step, or model.checkpoint to a full fine-tuning checkpoint step; both load through the training checkpointer. Requested sizes snap to the release resolution group.
Install
From a source checkout, install the diffusion media dependencies:
See the Installation Guide and Diffusion Training and Fine-Tuning Guide.