> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.hunyuan_image3.layers

Building blocks of HunyuanImage-3.0: attention, norms, timestep embedders and the UNet image projections.

Module and parameter names follow the released checkpoint so most weights load without renaming.

## Module Contents

### Classes

| Name                                                                                                       | Description                                                                                       |
| ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| [`GroupNorm32`](#nemo_automodel-components-models-hunyuan_image3-layers-GroupNorm32)                       | 32-group GroupNorm computed in fp32 (the reference runs it under autocast, which upcasts it).     |
| [`HunyuanImage3Attention`](#nemo_automodel-components-models-hunyuan_image3-layers-HunyuanImage3Attention) | GQA self-attention with a fused, KV-head-grouped QKV projection.                                  |
| [`HunyuanRMSNorm`](#nemo_automodel-components-models-hunyuan_image3-layers-HunyuanRMSNorm)                 | RMSNorm that normalizes in fp32 and returns `weight * x` in the input dtype.                      |
| [`HunyuanSharedMLP`](#nemo_automodel-components-models-hunyuan_image3-layers-HunyuanSharedMLP)             | SwiGLU MLP with the release's fused projection: `gate_and_up_proj` holds `[up; gate]` (up first). |
| [`ResBlock`](#nemo_automodel-components-models-hunyuan_image3-layers-ResBlock)                             | Residual conv block with timestep-conditioned adaptive GroupNorm (no up/down sampling).           |
| [`TimestepEmbedder`](#nemo_automodel-components-models-hunyuan_image3-layers-TimestepEmbedder)             | Sinusoidal timestep features followed by a two-layer GELU MLP.                                    |
| [`UNetDown`](#nemo_automodel-components-models-hunyuan_image3-layers-UNetDown)                             | Latent `[B, C, H, W]` -> token sequence `[B, H*W, hidden]` (patch size 1).                        |
| [`UNetUp`](#nemo_automodel-components-models-hunyuan_image3-layers-UNetUp)                                 | Token sequence `[B, H*W, hidden]` -> latent velocity `[B, C, H, W]` (patch size 1).               |

### Functions

| Name                                                                                               | Description                                                                                          |
| -------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| [`init_leaf_weights`](#nemo_automodel-components-models-hunyuan_image3-layers-init_leaf_weights)   | Default init for the non-MoE modules: linears and convs normal, norm affine = identity, biases zero. |
| [`timestep_embedding`](#nemo_automodel-components-models-hunyuan_image3-layers-timestep_embedding) | Sinusoidal timestep features, cosine half first.                                                     |

### API

```python
class nemo_automodel.components.models.hunyuan_image3.layers.GroupNorm32(
    channels: int,
    dtype: torch.dtype | None = None
)
```

**Bases:** `GroupNorm`

32-group GroupNorm computed in fp32 (the reference runs it under autocast, which upcasts it).

```python
nemo_automodel.components.models.hunyuan_image3.layers.GroupNorm32.forward(
    x: torch.Tensor
) -> torch.Tensor
```

Normalize channel groups.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[batch, channels, height, width].

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, channels, height, width] in the dtype of `x`.

```python
class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanImage3Attention(
    config: nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config,
    backend: nemo_automodel.components.models.common.BackendConfig,
    dtype: torch.dtype
)
```

**Bases:** `Module`

GQA self-attention with a fused, KV-head-grouped QKV projection.

`qkv_proj` outputs, for each KV head, its group of query heads followed by one key and one value head:
`[kv_heads, (q_per_kv + 2), head_dim]`. RoPE is applied before the per-head QK RMSNorm.

**`head_dim`** `= config.attention_head_dim`

---

**`key_layernorm`**

---

**`num_heads`** `= config.num_attention_heads`

---

**`num_kv_heads`** `= config.num_key_value_heads`

---

**`o_proj`**

---

**`q_per_kv`** `= self.num_heads // self.num_kv_heads`

---

**`qkv_proj`**

---

**`query_layernorm`**

---

**`use_qk_norm`** `= config.use_qk_norm`

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanImage3Attention.forward(
    x: torch.Tensor,
    cos: torch.Tensor,
    sin: torch.Tensor,
    attention_mask: torch.Tensor | None
) -> torch.Tensor
```

Attend over the joint sequence.

**Parameters:**

**`x`** `torch.Tensor`

`[batch, seq, hidden]` input.

---

**`cos`** `torch.Tensor`

`[batch, seq, head_dim]` fp32 rotary table.

---

**`sin`** `torch.Tensor`

`[batch, seq, head_dim]` fp32 rotary table.

---

**`attention_mask`** `torch.Tensor | None`

Boolean `[batch, 1, seq, seq]` mask, true where attention is allowed; `None` means
plain causal attention.

---

**Returns:** `torch.Tensor`

`[batch, seq, hidden]` attention output.

```python
class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm(
    hidden_size: int,
    eps: float = 1e-05,
    dtype: torch.dtype | None = None
)
```

**Bases:** `Module`

RMSNorm that normalizes in fp32 and returns `weight * x` in the input dtype.

**`weight`** `= nn.Parameter(torch.ones(hidden_size, dtype=dtype))`

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm.forward(
    x: torch.Tensor
) -> torch.Tensor
```

Normalize over the last axis.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[..., hidden], with arbitrary leading dimensions.

---

**Returns:** `torch.Tensor`

Tensor of shape \[..., hidden] in the promoted dtype of `weight` and `x`.

```python
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm.reset_parameters() -> None
```

```python
class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanSharedMLP(
    dim: int,
    inter_dim: int,
    backend: nemo_automodel.components.models.common.BackendConfig,
    bias: bool,
    dtype: torch.dtype
)
```

**Bases:** `Module`

SwiGLU MLP with the release's fused projection: `gate_and_up_proj` holds `[up; gate]` (up first).

Keeping the released layout lets the shared expert load and save by renaming alone.

**`down_proj`**

---

**`gate_and_up_proj`**

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanSharedMLP.forward(
    x: torch.Tensor
) -> torch.Tensor
```

Apply the shared expert.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[tokens, hidden].

---

**Returns:** `torch.Tensor`

Tensor of shape \[tokens, hidden].

```python
class nemo_automodel.components.models.hunyuan_image3.layers.ResBlock(
    in_channels: int,
    emb_channels: int,
    out_channels: int,
    dtype: torch.dtype | None = None
)
```

**Bases:** `Module`

Residual conv block with timestep-conditioned adaptive GroupNorm (no up/down sampling).

**`emb_layers`**

---

**`in_layers`**

---

**`out_layers`**

---

**`skip_connection`**

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.ResBlock.forward(
    x: torch.Tensor,
    emb: torch.Tensor
) -> torch.Tensor
```

Run the block.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[batch, in\_channels, height, width].

---

**`emb`** `torch.Tensor`

Tensor of shape \[batch, emb\_channels] timestep embedding.

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, out\_channels, height, width].

```python
class nemo_automodel.components.models.hunyuan_image3.layers.TimestepEmbedder(
    hidden_size: int,
    frequency_embedding_size: int = 256,
    dtype: torch.dtype | None = None
)
```

**Bases:** `Module`

Sinusoidal timestep features followed by a two-layer GELU MLP.

**`mlp`**

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.TimestepEmbedder.forward(
    t: torch.Tensor
) -> torch.Tensor
```

Embed timesteps.

**Parameters:**

**`t`** `torch.Tensor`

Tensor of shape \[batch] holding timesteps in \[0, 1000].

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, hidden] in the MLP weight dtype.

```python
class nemo_automodel.components.models.hunyuan_image3.layers.UNetDown(
    in_channels: int,
    emb_channels: int,
    hidden_channels: int,
    out_channels: int,
    dtype: torch.dtype | None
)
```

**Bases:** `Module`

Latent `[B, C, H, W]` -> token sequence `[B, H*W, hidden]` (patch size 1).

**`model`**

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.UNetDown.forward(
    x: torch.Tensor,
    emb: torch.Tensor
) -> torch.Tensor
```

Embed latents as tokens.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[batch, channels, height, width] VAE latents.

---

**`emb`** `torch.Tensor`

Tensor of shape \[batch, hidden] timestep embedding.

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, height \* width, hidden], tokens in row-major (height, width) order.

```python
class nemo_automodel.components.models.hunyuan_image3.layers.UNetUp(
    in_channels: int,
    emb_channels: int,
    hidden_channels: int,
    out_channels: int,
    dtype: torch.dtype | None
)
```

**Bases:** `Module`

Token sequence `[B, H*W, hidden]` -> latent velocity `[B, C, H, W]` (patch size 1).

**`model`**

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.UNetUp.forward(
    x: torch.Tensor,
    emb: torch.Tensor,
    token_h: int,
    token_w: int
) -> torch.Tensor
```

Project image tokens back to latent space.

**Parameters:**

**`x`** `torch.Tensor`

Tensor of shape \[batch, token\_h \* token\_w, hidden], tokens in row-major (height, width) order.

---

**`emb`** `torch.Tensor`

Tensor of shape \[batch, hidden] timestep embedding.

---

**`token_h`** `int`

Image height in tokens.

---

**`token_w`** `int`

Image width in tokens.

---

**Returns:** `torch.Tensor`

Tensor of shape \[batch, channels, token\_h, token\_w].

```python
nemo_automodel.components.models.hunyuan_image3.layers.init_leaf_weights(
    module: torch.nn.Module,
    init_std: float = 0.02
) -> None
```

Default init for the non-MoE modules: linears and convs normal, norm affine = identity, biases zero.

**Parameters:**

**`module`** `nn.Module`

Module tree whose leaves hold `weight` (and optionally `bias`) parameters: linears
(`[out, in]`), convs (`[out, in, kh, kw]`) and norms (`[hidden]`).

---

**`init_std`** `float` — default: 0.02

Standard deviation of the normal init for 2-D and 4-D weights.

---

```python
nemo_automodel.components.models.hunyuan_image3.layers.timestep_embedding(
    t: torch.Tensor,
    dim: int,
    max_period: float = 10000.0
) -> torch.Tensor
```

Sinusoidal timestep features, cosine half first.

**Parameters:**

**`t`** `torch.Tensor`

Tensor of shape \[batch] holding (possibly fractional) timesteps.

---

**`dim`** `int`

Feature size.

---

**`max_period`** `float` — default: 10000.0

Longest sinusoid period.

---

**Returns:** `torch.Tensor`

fp32 tensor of shape \[batch, dim].