> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.hunyuan_image3.model

HunyuanImage-3.0 (tencent/HunyuanImage-3.0) for flow-matching text-to-image training.

The released model is a single 80B-total / 13B-active MoE decoder that handles text and images in one sequence.
For text-to-image generation the sequence is::

\<bos> prompt \<boi> \<img\_size\_*> \<img\_ratio\_*> \<timestep> \<img> \* (h\*w) \<eoi>

The `&lt;timestep&gt;` slot receives a timestep embedding and the `&lt;img&gt;` slots receive the noisy VAE latents through
a UNet patch embedding. Text tokens attend causally; the image tokens attend to each other bidirectionally. The
final hidden states at the image slots (before the final norm) go through a UNet final layer that predicts the
flow velocity `noise - x0`.

The decoder uses Automodel's `MoE` under `model.layers[*].mlp` so
the MoE parallelizer can apply FSDP2 and expert parallelism. The VAE and the vision encoder of the release are not
part of this module: latents are precomputed during preprocessing, and image conditioning is not supported.

## Module Contents

### Classes

| Name                                                                                                          | Description                                                                                                |
| ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| [`HunyuanImage3Block`](#nemo_automodel-components-models-hunyuan_image3-model-HunyuanImage3Block)             | Pre-norm decoder layer: GQA attention, then routed experts (`MoE`) plus the release's fused shared expert. |
| [`HunyuanImage3ForCausalMM`](#nemo_automodel-components-models-hunyuan_image3-model-HunyuanImage3ForCausalMM) | HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.                                 |
| [`HunyuanImage3Model`](#nemo_automodel-components-models-hunyuan_image3-model-HunyuanImage3Model)             | Decoder backbone: token embedding, MoE layers and the final norm (`wte` / `ln_f` in the release).          |

### Functions

| Name                                                                                                              | Description                                                                                         |
| ----------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| [`_model_dtype`](#nemo_automodel-components-models-hunyuan_image3-model-_model_dtype)                             | Parameter dtype from the config (`dtype` under transformers 5, `torch_dtype` before).               |
| [`build_joint_attention_mask`](#nemo_automodel-components-models-hunyuan_image3-model-build_joint_attention_mask) | Boolean `[batch, 1, seq, seq]` mask: causal text, bidirectional image span, padded keys masked out. |
| [`build_moe_config`](#nemo_automodel-components-models-hunyuan_image3-model-build_moe_config)                     | MoE settings of the release: softmax over all experts, top-k, renormalized, one shared expert.      |

### Data

[`ModelClass`](#nemo_automodel-components-models-hunyuan_image3-model-ModelClass)

### API

```python
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig,
    backend: nemo_automodel.components.models.common.BackendConfig,
    dtype: torch.dtype
)
```

**Bases:** `Module`

Pre-norm decoder layer: GQA attention, then routed experts (`MoE`) plus the release's fused shared expert.

The shared expert lives in the block, not in `MoE`, because its fused `[up; gate]` projection is not the
`MLP` that `MoE` builds and manages.

**`input_layernorm`**

---

**`mlp`** `= MoE(moe_config, backend)`

---

**`post_attention_layernorm`**

---

**`self_attn`** `= HunyuanImage3Attention(config, backend, dtype)`

---

**`shared_mlp`**

---

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block.forward(
    x: torch.Tensor,
    cos: torch.Tensor,
    sin: torch.Tensor,
    attention_mask: torch.Tensor | None,
    padding_mask: torch.Tensor | None
) -> torch.Tensor
```

Run one decoder layer.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[batch, sequence, hidden].

---

**`cos`** `torch.Tensor`

fp32 tensor of shape \[batch, sequence, head\_dim] rotary table.

---

**`sin`** `torch.Tensor`

fp32 tensor of shape \[batch, sequence, head\_dim] rotary table.

---

**`attention_mask`** `torch.Tensor | None`

Boolean tensor of shape \[batch, 1, sequence, sequence], true where attention is
allowed, or `None` for causal attention.

---

**`padding_mask`** `torch.Tensor | None`

Boolean tensor of shape \[batch, sequence], true at padding, or `None`.

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, sequence, hidden].

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block.init_weights(
    buffer_device: torch.device,
    init_std: float = 0.02
) -> None
```

```python
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig | None = None,
    backend: nemo_automodel.components.models.common.BackendConfig | None = None,
    kwargs: typing.Any = {}
)
```

**Bases:** [HFCheckpointingMixin](/nemo-automodel/nemo_automodel/components/models/common/hf_checkpointing_mixin#nemo_automodel-components-models-common-hf_checkpointing_mixin-HFCheckpointingMixin), `Module`, [MoEFSDPSyncMixin](/nemo-automodel/nemo_automodel/components/moe/fsdp_mixin#nemo_automodel-components-moe-fsdp_mixin-MoEFSDPSyncMixin)

HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.

`forward` predicts the flow velocity of noisy latents given the token sequence; `forward_text` returns
next-token logits (the release's `gen_text` mode) and is used for parity checks.

**`_keep_in_fp32_modules_strict`** `= ['mlp.gate']`

---

**`backend`** `= backend or BackendConfig()`

---

**`final_layer`**

---

**`lm_head`**

---

**`model`**

---

**`patch_embed`**

---

**`state_dict_adapter`**

---

**`tie_word_embeddings_support`** `TieSupport = TieSupport.UNTIED_ONLY`

---

**`time_embed`** `= TimestepEmbedder(hidden, dtype=dtype)`

---

**`time_embed_2`** `= TimestepEmbedder(hidden, dtype=dtype)`

---

**`timestep_emb`** `= TimestepEmbedder(hidden, dtype=dtype)`

---

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward(
    input_ids: torch.Tensor,
    latents: torch.Tensor,
    timestep: torch.Tensor,
    valid_lengths: torch.Tensor | None = None,
    return_dict: bool = False,
    kwargs: typing.Any = {}
) -> tuple[torch.Tensor] | dict[str, torch.Tensor]
```

Predict the flow velocity of the noisy latents.

**Parameters:**

**`input_ids`** `torch.Tensor`

`[batch, seq]` token ids laid out as described in the module docstring, right-padded.
Every row holds exactly `h*w` contiguous image tokens, preceded by the `&lt;timestep&gt;` token.

---

**`latents`** `torch.Tensor`

`[batch, channels, h, w]` noisy VAE latents.

---

**`timestep`** `torch.Tensor`

`[batch]` flow-matching timesteps in `[0, 1000]` (`sigma * 1000`).

---

**`valid_lengths`** `torch.Tensor | None` — default: None

`[batch]` number of non-padding tokens per row; defaults to the full length.

---

**`return_dict`** `bool` — default: False

Return `&#123;"sample": velocity&#125;` instead of `(velocity,)`.

---

**Returns:** `tuple[torch.Tensor] | dict[str, torch.Tensor]`

`[batch, channels, h, w]` predicted velocity, as a one-element tuple or a dict.

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward_text(
    input_ids: torch.Tensor
) -> torch.Tensor
```

Next-token logits of a causal text-only sequence (the release's `gen_text` mode).

**Parameters:**

**`input_ids`** `torch.Tensor`

Long tensor of shape \[batch, sequence].

---

**Returns:** `torch.Tensor`

fp32 tensor of shape \[batch, sequence, vocab].

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_config(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig | None = None,
    backend: nemo_automodel.components.models.common.BackendConfig | None = None,
    kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
```

classmethod

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_pretrained(
    pretrained_model_name_or_path: str,
    model_args: typing.Any = (),
    kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
```

classmethod

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_input_embeddings() -> torch.nn.Embedding
```

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_output_embeddings() -> torch.nn.Module
```

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.initialize_weights(
    buffer_device: torch.device | None = None,
    dtype: torch.dtype = torch.bfloat16
) -> None
```

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.update_moe_gate_bias() -> None
```

```python
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    backend: nemo_automodel.components.models.common.BackendConfig,
    moe_config: nemo_automodel.components.moe.config.MoEConfig,
    dtype: torch.dtype
)
```

**Bases:** `Module`

Decoder backbone: token embedding, MoE layers and the final norm (`wte` / `ln_f` in the release).

**`embed_tokens`**

---

**`layers`**

---

**`norm`**

---

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model.forward(
    inputs_embeds: torch.Tensor,
    cos: torch.Tensor,
    sin: torch.Tensor,
    attention_mask: torch.Tensor | None = None,
    padding_mask: torch.Tensor | None = None
) -> torch.Tensor
```

Run the decoder layers.

**Parameters:**

**`inputs_embeds`** `torch.Tensor`

Tensor of shape \[batch, sequence, hidden].

---

**`cos`** `torch.Tensor`

fp32 tensor of shape \[batch, sequence, head\_dim] rotary table.

---

**`sin`** `torch.Tensor`

fp32 tensor of shape \[batch, sequence, head\_dim] rotary table.

---

**`attention_mask`** `torch.Tensor | None` — default: None

Boolean tensor of shape \[batch, 1, sequence, sequence], true where attention is
allowed, or `None` for causal attention.

---

**`padding_mask`** `torch.Tensor | None` — default: None

Boolean tensor of shape \[batch, sequence], true at padding, or `None`.

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, sequence, hidden]: the last hidden states *before* the final norm.

```python
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model.init_weights(
    buffer_device: torch.device,
    init_std: float = 0.02
) -> None
```

```python
nemo_automodel.components.models.hunyuan_image3.model._model_dtype(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config
) -> torch.dtype
```

Parameter dtype from the config (`dtype` under transformers 5, `torch_dtype` before).

```python
nemo_automodel.components.models.hunyuan_image3.model.build_joint_attention_mask(
    seq_len: int,
    image_starts: torch.Tensor,
    num_image_tokens: int,
    valid_lengths: torch.Tensor
) -> torch.Tensor
```

Boolean `[batch, 1, seq, seq]` mask: causal text, bidirectional image span, padded keys masked out.

**Parameters:**

**`seq_len`** `int`

Padded sequence length.

---

**`image_starts`** `torch.Tensor`

Long tensor of shape \[batch], index of the first image token of every sample.

---

**`num_image_tokens`** `int`

Number of image tokens (same for every sample of a batch).

---

**`valid_lengths`** `torch.Tensor`

Long tensor of shape \[batch], number of real (non-padding) tokens of every sample.

---

**Returns:** `torch.Tensor`

Boolean tensor of shape \[batch, 1, sequence, sequence], indexed \[batch, 1, query, key], true where

```python
nemo_automodel.components.models.hunyuan_image3.model.build_moe_config(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    overrides: dict[str, typing.Any] | None = None
) -> nemo_automodel.components.moe.config.MoEConfig
```

MoE settings of the release: softmax over all experts, top-k, renormalized, one shared expert.

```python
nemo_automodel.components.models.hunyuan_image3.model.ModelClass = HunyuanImage3ForCausalMM
```