> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# HunyuanImage-3.0

> Fine-tune tencent/HunyuanImage-3.0, an 80B MoE text-to-image transformer, with LoRA and expert parallelism in NeMo AutoModel.

[HunyuanImage-3.0](https://huggingface.co/tencent/HunyuanImage-3.0) is Tencent's text-to-image model. Unlike the diffusers-based models on this site, it is a single multimodal MoE decoder: the prompt tokens and the image tokens share one sequence, and the same transformer predicts the flow-matching velocity of the image latents. NeMo AutoModel ships its own implementation of the transformer, so the diffusion recipe trains it through the custom-model path with FSDP2 and expert parallelism.

## Model Reference

### Model Architecture

| Property         | Value                                                                                                                  |
| ---------------- | ---------------------------------------------------------------------------------------------------------------------- |
| **Tasks**        | Text-to-Image                                                                                                          |
| **Architecture** | Multimodal MoE decoder with flow-matching image head (80B total / 13B active)                                          |
| **Layers**       | 32, each with GQA attention (32 query / 8 KV heads, head dim 128) and 64 routed experts (top-8) plus one shared expert |
| **Positions**    | 2D RoPE over the joint text / image sequence                                                                           |
| **Latents**      | 32 channels, 16x spatial compression (release VAE)                                                                     |
| **HF Org**       | [tencent](https://huggingface.co/tencent)                                                                              |

### Model Class

`HunyuanImage3ForCausalMM` (`nemo_automodel/components/models/hunyuan_image3/`) is registered for the checkpoint's `HunyuanImage3ForCausalMM` architecture and `hunyuan_image_3_moe` model type, so the config loads without `trust_remote_code`. The decoder uses the shared `MoE` layer, which the MoE parallelizer shards with FSDP2 and expert parallelism; the released checkpoint is converted by `HunyuanImage3StateDictAdapter` (fused `[up; gate]` expert projections to grouped experts, `wte` / `ln_f` renames). The release's VAE and vision encoder are not part of the training model.

Supported: FSDP2, expert parallelism, LoRA, full fine-tuning through the same recipe. Not supported: tensor, pipeline and context parallelism, image conditioning (image-to-image), text generation training.

## Available Models

| Model            | HF ID                                                                         | Training Scope                          |
| ---------------- | ----------------------------------------------------------------------------- | --------------------------------------- |
| HunyuanImage-3.0 | [`tencent/HunyuanImage-3.0`](https://huggingface.co/tencent/HunyuanImage-3.0) | Text-to-image LoRA and full fine-tuning |

## Recipes

| Recipe                                                                                                                                                    | Description                                                                      |
| --------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| [hunyuan\_image3\_t2i\_flow\_lora.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/diffusion/finetune/hunyuan_image3_t2i_flow_lora.yaml) | Single-node LoRA fine-tuning with 8-way expert parallelism                       |
| [hunyuan\_image3\_t2i\_flow.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/diffusion/finetune/hunyuan_image3_t2i_flow.yaml)            | Four-node full fine-tuning with 8-way expert parallelism                         |
| [generate\_hunyuan\_image3.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/diffusion/generate/configs/generate_hunyuan_image3.yaml)     | Text-to-image generation from the base model, a full fine-tune or a LoRA adapter |

## Prepare Data

The preprocessing step needs the release's VAE and tokenizer wrapper, which are loaded from the checkpoint's remote code (`nemo_automodel/components/models/hunyuan_image3/release.py`, shared with generation). For every image it stores the VAE latent and the release's prompt token ids around the image span (the conditional prompt, the `<cfg>` unconditional prompt used for classifier-free guidance dropout, and the suffix). Images are snapped to the release's own resolution group (33 aspect ratios around 1024x1024) instead of the generic buckets:

```bash
uv run python -m tools.diffusion.preprocessing_multiprocess image \
  --image_dir /data/images \
  --output_dir /cache/hunyuan-image3 \
  --processor hunyuan_image3 \
  --model_name tencent/HunyuanImage-3.0 \
  --resolution_preset 1024p
```

Captions follow the image processors' JSONL convention (`<stem>_internvl.json` next to the image, or the `--caption_field` you pass).

## Cache Contract

Each cache sample contains:

* `latent`: `[32, height / 16, width / 16]` (fp32, scaled by the VAE scaling factor)
* `prompt_input_ids`: token ids before the image span, ending in `<timestep>`
* `uncond_prompt_input_ids`: the same with the prompt replaced by `<cfg>` tokens
* `prompt_suffix_ids`: token ids after the image span (`<eoi>`)
* `bucket_resolution`, `original_resolution`, `crop_offset`, `prompt`, `image_path`

The `hunyuan_image3` flow-matching adapter assembles `prefix + <img> * (h * w) + suffix` per sample and right-pads the batch; `flow_matching.cfg_dropout_prob` swaps in the unconditional prefix.

## Train

Set `data.dataloader.cache_dir` and `checkpoint.checkpoint_dir`, then launch one rank per GPU. LoRA fits on one node:

```bash
uv run torchrun --nproc-per-node=8 \
  examples/diffusion/finetune/finetune.py \
  -c examples/diffusion/finetune/hunyuan_image3_t2i_flow_lora.yaml
```

Full fine-tuning keeps fp32-precision master weights and Adam moments for all 80B parameters (about 1.1 TB), so its recipe is laid out for four nodes; run the same command on every node with its own `--node-rank`:

```bash
uv run torchrun --nnodes=4 --node-rank=NODE_RANK --nproc-per-node=8 \
  --master-addr=MASTER_ADDR --master-port=29500 \
  examples/diffusion/finetune/finetune.py \
  -c examples/diffusion/finetune/hunyuan_image3_t2i_flow.yaml
```

Both recipes keep the router in fp32 (`gate_precision: float32`), train with `flow_shift: 3.0` (the value the release samples with) and use Transformer Engine's `FusedAdam` with fp32 master weights (`master_weights: true` with 16-bit remainders), since plain AdamW on bf16 parameters rounds away most small updates. The LoRA recipe adapts the attention projections and the shared expert, and its checkpoints hold only the adapter (`adapter_model.safetensors` plus its config). `fsdp.ep_size: 1` is also supported; the experts are then sharded by FSDP2.

## Generate

`examples/diffusion/generate/generate.py` samples with the native transformer, sharded across the launched ranks as in training, and the release's VAE, prompt format and `flow_shift` from the checkpoint. Sampling follows the release pipeline: Euler steps on the shifted sigma schedule with classifier-free guidance over a `[prompt, <cfg>]` batch. Every rank samples the same prompts:

```bash
uv run torchrun --nproc-per-node=8 \
  examples/diffusion/generate/generate.py \
  -c examples/diffusion/generate/configs/generate_hunyuan_image3.yaml
```

Set `model.lora_weights` to the `model/` directory of a LoRA checkpoint step, or `model.checkpoint` to a full fine-tuning checkpoint step; both load through the training checkpointer. Requested sizes snap to the release resolution group.

## Install

From a source checkout, install the diffusion media dependencies:

```bash
uv sync --locked --all-groups --extra all --extra diffusion-media
```

See the [Installation Guide](/get-started/installation) and [Diffusion Training and Fine-Tuning Guide](/recipes-e2e-examples/diffusion-fine-tuning).