nemo_automodel.components.models.hunyuan_image3.layers

View as Markdown

Building blocks of HunyuanImage-3.0: attention, norms, timestep embedders and the UNet image projections.

Module and parameter names follow the released checkpoint so most weights load without renaming.

Module Contents

Classes

NameDescription
GroupNorm3232-group GroupNorm computed in fp32 (the reference runs it under autocast, which upcasts it).
HunyuanImage3AttentionGQA self-attention with a fused, KV-head-grouped QKV projection.
HunyuanRMSNormRMSNorm that normalizes in fp32 and returns weight * x in the input dtype.
HunyuanSharedMLPSwiGLU MLP with the release’s fused projection: gate_and_up_proj holds [up; gate] (up first).
ResBlockResidual conv block with timestep-conditioned adaptive GroupNorm (no up/down sampling).
TimestepEmbedderSinusoidal timestep features followed by a two-layer GELU MLP.
UNetDownLatent [B, C, H, W] -> token sequence [B, H*W, hidden] (patch size 1).
UNetUpToken sequence [B, H*W, hidden] -> latent velocity [B, C, H, W] (patch size 1).

Functions

NameDescription
init_leaf_weightsDefault init for the non-MoE modules: linears and convs normal, norm affine = identity, biases zero.
timestep_embeddingSinusoidal timestep features, cosine half first.

API

class nemo_automodel.components.models.hunyuan_image3.layers.GroupNorm32(
channels: int,
dtype: torch.dtype | None = None
)

Bases: GroupNorm

32-group GroupNorm computed in fp32 (the reference runs it under autocast, which upcasts it).

nemo_automodel.components.models.hunyuan_image3.layers.GroupNorm32.forward(
x: torch.Tensor
) -> torch.Tensor

Normalize channel groups.

Parameters:

x
torch.Tensor

Tensor of shape [batch, channels, height, width].

Returns: torch.Tensor

Tensor of shape [batch, channels, height, width] in the dtype of x.

class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanImage3Attention(
dtype: torch.dtype
)

Bases: Module

GQA self-attention with a fused, KV-head-grouped QKV projection.

qkv_proj outputs, for each KV head, its group of query heads followed by one key and one value head: [kv_heads, (q_per_kv + 2), head_dim]. RoPE is applied before the per-head QK RMSNorm.

head_dim
= config.attention_head_dim
key_layernorm
num_heads
= config.num_attention_heads
num_kv_heads
= config.num_key_value_heads
o_proj
q_per_kv
= self.num_heads // self.num_kv_heads
qkv_proj
query_layernorm
use_qk_norm
= config.use_qk_norm
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanImage3Attention.forward(
x: torch.Tensor,
cos: torch.Tensor,
sin: torch.Tensor,
attention_mask: torch.Tensor | None
) -> torch.Tensor

Attend over the joint sequence.

Parameters:

x
torch.Tensor

[batch, seq, hidden] input.

cos
torch.Tensor

[batch, seq, head_dim] fp32 rotary table.

sin
torch.Tensor

[batch, seq, head_dim] fp32 rotary table.

attention_mask
torch.Tensor | None

Boolean [batch, 1, seq, seq] mask, true where attention is allowed; None means plain causal attention.

Returns: torch.Tensor

[batch, seq, hidden] attention output.

class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm(
hidden_size: int,
eps: float = 1e-05,
dtype: torch.dtype | None = None
)

Bases: Module

RMSNorm that normalizes in fp32 and returns weight * x in the input dtype.

weight
= nn.Parameter(torch.ones(hidden_size, dtype=dtype))
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm.forward(
x: torch.Tensor
) -> torch.Tensor

Normalize over the last axis.

Parameters:

x
torch.Tensor

Tensor of shape […, hidden], with arbitrary leading dimensions.

Returns: torch.Tensor

Tensor of shape […, hidden] in the promoted dtype of weight and x.

nemo_automodel.components.models.hunyuan_image3.layers.HunyuanRMSNorm.reset_parameters() -> None
class nemo_automodel.components.models.hunyuan_image3.layers.HunyuanSharedMLP(
dim: int,
inter_dim: int,
bias: bool,
dtype: torch.dtype
)

Bases: Module

SwiGLU MLP with the release’s fused projection: gate_and_up_proj holds [up; gate] (up first).

Keeping the released layout lets the shared expert load and save by renaming alone.

down_proj
gate_and_up_proj
nemo_automodel.components.models.hunyuan_image3.layers.HunyuanSharedMLP.forward(
x: torch.Tensor
) -> torch.Tensor

Apply the shared expert.

Parameters:

x
torch.Tensor

Tensor of shape [tokens, hidden].

Returns: torch.Tensor

Tensor of shape [tokens, hidden].

class nemo_automodel.components.models.hunyuan_image3.layers.ResBlock(
in_channels: int,
emb_channels: int,
out_channels: int,
dtype: torch.dtype | None = None
)

Bases: Module

Residual conv block with timestep-conditioned adaptive GroupNorm (no up/down sampling).

emb_layers
in_layers
out_layers
skip_connection
nemo_automodel.components.models.hunyuan_image3.layers.ResBlock.forward(
x: torch.Tensor,
emb: torch.Tensor
) -> torch.Tensor

Run the block.

Parameters:

x
torch.Tensor

Tensor of shape [batch, in_channels, height, width].

emb
torch.Tensor

Tensor of shape [batch, emb_channels] timestep embedding.

Returns: torch.Tensor

Tensor of shape [batch, out_channels, height, width].

class nemo_automodel.components.models.hunyuan_image3.layers.TimestepEmbedder(
hidden_size: int,
frequency_embedding_size: int = 256,
dtype: torch.dtype | None = None
)

Bases: Module

Sinusoidal timestep features followed by a two-layer GELU MLP.

mlp
nemo_automodel.components.models.hunyuan_image3.layers.TimestepEmbedder.forward(
t: torch.Tensor
) -> torch.Tensor

Embed timesteps.

Parameters:

t
torch.Tensor

Tensor of shape [batch] holding timesteps in [0, 1000].

Returns: torch.Tensor

Tensor of shape [batch, hidden] in the MLP weight dtype.

class nemo_automodel.components.models.hunyuan_image3.layers.UNetDown(
in_channels: int,
emb_channels: int,
hidden_channels: int,
out_channels: int,
dtype: torch.dtype | None
)

Bases: Module

Latent [B, C, H, W] -> token sequence [B, H*W, hidden] (patch size 1).

model
nemo_automodel.components.models.hunyuan_image3.layers.UNetDown.forward(
x: torch.Tensor,
emb: torch.Tensor
) -> torch.Tensor

Embed latents as tokens.

Parameters:

x
torch.Tensor

Tensor of shape [batch, channels, height, width] VAE latents.

emb
torch.Tensor

Tensor of shape [batch, hidden] timestep embedding.

Returns: torch.Tensor

Tensor of shape [batch, height * width, hidden], tokens in row-major (height, width) order.

class nemo_automodel.components.models.hunyuan_image3.layers.UNetUp(
in_channels: int,
emb_channels: int,
hidden_channels: int,
out_channels: int,
dtype: torch.dtype | None
)

Bases: Module

Token sequence [B, H*W, hidden] -> latent velocity [B, C, H, W] (patch size 1).

model
nemo_automodel.components.models.hunyuan_image3.layers.UNetUp.forward(
x: torch.Tensor,
emb: torch.Tensor,
token_h: int,
token_w: int
) -> torch.Tensor

Project image tokens back to latent space.

Parameters:

x
torch.Tensor

Tensor of shape [batch, token_h * token_w, hidden], tokens in row-major (height, width) order.

emb
torch.Tensor

Tensor of shape [batch, hidden] timestep embedding.

token_h
int

Image height in tokens.

token_w
int

Image width in tokens.

Returns: torch.Tensor

Tensor of shape [batch, channels, token_h, token_w].

nemo_automodel.components.models.hunyuan_image3.layers.init_leaf_weights(
module: torch.nn.Module,
init_std: float = 0.02
) -> None

Default init for the non-MoE modules: linears and convs normal, norm affine = identity, biases zero.

Parameters:

module
nn.Module

Module tree whose leaves hold weight (and optionally bias) parameters: linears ([out, in]), convs ([out, in, kh, kw]) and norms ([hidden]).

init_std
floatDefaults to 0.02

Standard deviation of the normal init for 2-D and 4-D weights.

nemo_automodel.components.models.hunyuan_image3.layers.timestep_embedding(
t: torch.Tensor,
dim: int,
max_period: float = 10000.0
) -> torch.Tensor

Sinusoidal timestep features, cosine half first.

Parameters:

t
torch.Tensor

Tensor of shape [batch] holding (possibly fractional) timesteps.

dim
int

Feature size.

max_period
floatDefaults to 10000.0

Longest sinusoid period.

Returns: torch.Tensor

fp32 tensor of shape [batch, dim].