nemo_automodel.components.models.hunyuan_image3.layers
nemo_automodel.components.models.hunyuan_image3.layers
Building blocks of HunyuanImage-3.0: attention, norms, timestep embedders and the UNet image projections.
Module and parameter names follow the released checkpoint so most weights load without renaming.
Module Contents
Classes
Functions
API
Bases: GroupNorm
32-group GroupNorm computed in fp32 (the reference runs it under autocast, which upcasts it).
Normalize channel groups.
Parameters:
Tensor of shape [batch, channels, height, width].
Returns: torch.Tensor
Tensor of shape [batch, channels, height, width] in the dtype of x.
Bases: Module
GQA self-attention with a fused, KV-head-grouped QKV projection.
qkv_proj outputs, for each KV head, its group of query heads followed by one key and one value head:
[kv_heads, (q_per_kv + 2), head_dim]. RoPE is applied before the per-head QK RMSNorm.
Attend over the joint sequence.
Parameters:
[batch, seq, hidden] input.
[batch, seq, head_dim] fp32 rotary table.
[batch, seq, head_dim] fp32 rotary table.
Boolean [batch, 1, seq, seq] mask, true where attention is allowed; None means
plain causal attention.
Returns: torch.Tensor
[batch, seq, hidden] attention output.
Bases: Module
RMSNorm that normalizes in fp32 and returns weight * x in the input dtype.
Normalize over the last axis.
Parameters:
Tensor of shape […, hidden], with arbitrary leading dimensions.
Returns: torch.Tensor
Tensor of shape […, hidden] in the promoted dtype of weight and x.
Bases: Module
SwiGLU MLP with the release’s fused projection: gate_and_up_proj holds [up; gate] (up first).
Keeping the released layout lets the shared expert load and save by renaming alone.
Apply the shared expert.
Parameters:
Tensor of shape [tokens, hidden].
Returns: torch.Tensor
Tensor of shape [tokens, hidden].
Bases: Module
Residual conv block with timestep-conditioned adaptive GroupNorm (no up/down sampling).
Run the block.
Parameters:
Tensor of shape [batch, in_channels, height, width].
Tensor of shape [batch, emb_channels] timestep embedding.
Returns: torch.Tensor
Tensor of shape [batch, out_channels, height, width].
Bases: Module
Sinusoidal timestep features followed by a two-layer GELU MLP.
Embed timesteps.
Parameters:
Tensor of shape [batch] holding timesteps in [0, 1000].
Returns: torch.Tensor
Tensor of shape [batch, hidden] in the MLP weight dtype.
Bases: Module
Latent [B, C, H, W] -> token sequence [B, H*W, hidden] (patch size 1).
Embed latents as tokens.
Parameters:
Tensor of shape [batch, channels, height, width] VAE latents.
Tensor of shape [batch, hidden] timestep embedding.
Returns: torch.Tensor
Tensor of shape [batch, height * width, hidden], tokens in row-major (height, width) order.
Bases: Module
Token sequence [B, H*W, hidden] -> latent velocity [B, C, H, W] (patch size 1).
Project image tokens back to latent space.
Parameters:
Tensor of shape [batch, token_h * token_w, hidden], tokens in row-major (height, width) order.
Tensor of shape [batch, hidden] timestep embedding.
Image height in tokens.
Image width in tokens.
Returns: torch.Tensor
Tensor of shape [batch, channels, token_h, token_w].
Default init for the non-MoE modules: linears and convs normal, norm affine = identity, biases zero.
Parameters:
Module tree whose leaves hold weight (and optionally bias) parameters: linears
([out, in]), convs ([out, in, kh, kw]) and norms ([hidden]).
Standard deviation of the normal init for 2-D and 4-D weights.
Sinusoidal timestep features, cosine half first.
Parameters:
Tensor of shape [batch] holding (possibly fractional) timesteps.
Feature size.
Longest sinusoid period.
Returns: torch.Tensor
fp32 tensor of shape [batch, dim].