nemo_automodel.components.models.hunyuan_image3.model

View as Markdown

HunyuanImage-3.0 (tencent/HunyuanImage-3.0) for flow-matching text-to-image training.

The released model is a single 80B-total / 13B-active MoE decoder that handles text and images in one sequence. For text-to-image generation the sequence is::

<bos> prompt <boi> <img_size_> <img_ratio_> <timestep> <img> * (h*w) <eoi>

The &lt;timestep&gt; slot receives a timestep embedding and the &lt;img&gt; slots receive the noisy VAE latents through a UNet patch embedding. Text tokens attend causally; the image tokens attend to each other bidirectionally. The final hidden states at the image slots (before the final norm) go through a UNet final layer that predicts the flow velocity noise - x0.

The decoder uses Automodel’s MoE under model.layers[*].mlp so the MoE parallelizer can apply FSDP2 and expert parallelism. The VAE and the vision encoder of the release are not part of this module: latents are precomputed during preprocessing, and image conditioning is not supported.

Module Contents

Classes

NameDescription
HunyuanImage3BlockPre-norm decoder layer: GQA attention, then routed experts (MoE) plus the release’s fused shared expert.
HunyuanImage3ForCausalMMHunyuanImage-3.0 transformer with the text-to-image flow-matching forward.
HunyuanImage3ModelDecoder backbone: token embedding, MoE layers and the final norm (wte / ln_f in the release).

Functions

NameDescription
_model_dtypeParameter dtype from the config (dtype under transformers 5, torch_dtype before).
build_joint_attention_maskBoolean [batch, 1, seq, seq] mask: causal text, bidirectional image span, padded keys masked out.
build_moe_configMoE settings of the release: softmax over all experts, top-k, renormalized, one shared expert.

Data

ModelClass

API

class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block(
dtype: torch.dtype
)

Bases: Module

Pre-norm decoder layer: GQA attention, then routed experts (MoE) plus the release’s fused shared expert.

The shared expert lives in the block, not in MoE, because its fused [up; gate] projection is not the MLP that MoE builds and manages.

input_layernorm
mlp
= MoE(moe_config, backend)
post_attention_layernorm
self_attn
= HunyuanImage3Attention(config, backend, dtype)
shared_mlp
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block.forward(
x: torch.Tensor,
cos: torch.Tensor,
sin: torch.Tensor,
attention_mask: torch.Tensor | None,
padding_mask: torch.Tensor | None
) -> torch.Tensor

Run one decoder layer.

Parameters:

x
torch.Tensor

Tensor of shape [batch, sequence, hidden].

cos
torch.Tensor

fp32 tensor of shape [batch, sequence, head_dim] rotary table.

sin
torch.Tensor

fp32 tensor of shape [batch, sequence, head_dim] rotary table.

attention_mask
torch.Tensor | None

Boolean tensor of shape [batch, 1, sequence, sequence], true where attention is allowed, or None for causal attention.

padding_mask
torch.Tensor | None

Boolean tensor of shape [batch, sequence], true at padding, or None.

Returns: torch.Tensor

Tensor of shape [batch, sequence, hidden].

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Block.init_weights(
buffer_device: torch.device,
init_std: float = 0.02
) -> None
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM(
kwargs: typing.Any = {}
)

Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin

HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.

forward predicts the flow velocity of noisy latents given the token sequence; forward_text returns next-token logits (the release’s gen_text mode) and is used for parity checks.

_keep_in_fp32_modules_strict
= ['mlp.gate']
backend
= backend or BackendConfig()
final_layer
lm_head
model
patch_embed
state_dict_adapter
tie_word_embeddings_support
TieSupport = TieSupport.UNTIED_ONLY
time_embed
= TimestepEmbedder(hidden, dtype=dtype)
time_embed_2
= TimestepEmbedder(hidden, dtype=dtype)
timestep_emb
= TimestepEmbedder(hidden, dtype=dtype)
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward(
input_ids: torch.Tensor,
latents: torch.Tensor,
timestep: torch.Tensor,
valid_lengths: torch.Tensor | None = None,
return_dict: bool = False,
kwargs: typing.Any = {}
) -> tuple[torch.Tensor] | dict[str, torch.Tensor]

Predict the flow velocity of the noisy latents.

Parameters:

input_ids
torch.Tensor

[batch, seq] token ids laid out as described in the module docstring, right-padded. Every row holds exactly h*w contiguous image tokens, preceded by the &lt;timestep&gt; token.

latents
torch.Tensor

[batch, channels, h, w] noisy VAE latents.

timestep
torch.Tensor

[batch] flow-matching timesteps in [0, 1000] (sigma * 1000).

valid_lengths
torch.Tensor | NoneDefaults to None

[batch] number of non-padding tokens per row; defaults to the full length.

return_dict
boolDefaults to False

Return &#123;"sample": velocity&#125; instead of (velocity,).

Returns: tuple[torch.Tensor] | dict[str, torch.Tensor]

[batch, channels, h, w] predicted velocity, as a one-element tuple or a dict.

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward_text(
input_ids: torch.Tensor
) -> torch.Tensor

Next-token logits of a causal text-only sequence (the release’s gen_text mode).

Parameters:

input_ids
torch.Tensor

Long tensor of shape [batch, sequence].

Returns: torch.Tensor

fp32 tensor of shape [batch, sequence, vocab].

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_config(
kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
classmethod
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_pretrained(
pretrained_model_name_or_path: str,
model_args: typing.Any = (),
kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
classmethod
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_input_embeddings() -> torch.nn.Embedding
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_output_embeddings() -> torch.nn.Module
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.initialize_weights(
buffer_device: torch.device | None = None,
dtype: torch.dtype = torch.bfloat16
) -> None
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.update_moe_gate_bias() -> None
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model(
dtype: torch.dtype
)

Bases: Module

Decoder backbone: token embedding, MoE layers and the final norm (wte / ln_f in the release).

embed_tokens
layers
norm
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model.forward(
inputs_embeds: torch.Tensor,
cos: torch.Tensor,
sin: torch.Tensor,
attention_mask: torch.Tensor | None = None,
padding_mask: torch.Tensor | None = None
) -> torch.Tensor

Run the decoder layers.

Parameters:

inputs_embeds
torch.Tensor

Tensor of shape [batch, sequence, hidden].

cos
torch.Tensor

fp32 tensor of shape [batch, sequence, head_dim] rotary table.

sin
torch.Tensor

fp32 tensor of shape [batch, sequence, head_dim] rotary table.

attention_mask
torch.Tensor | NoneDefaults to None

Boolean tensor of shape [batch, 1, sequence, sequence], true where attention is allowed, or None for causal attention.

padding_mask
torch.Tensor | NoneDefaults to None

Boolean tensor of shape [batch, sequence], true at padding, or None.

Returns: torch.Tensor

Tensor of shape [batch, sequence, hidden]: the last hidden states before the final norm.

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3Model.init_weights(
buffer_device: torch.device,
init_std: float = 0.02
) -> None
nemo_automodel.components.models.hunyuan_image3.model._model_dtype(
) -> torch.dtype

Parameter dtype from the config (dtype under transformers 5, torch_dtype before).

nemo_automodel.components.models.hunyuan_image3.model.build_joint_attention_mask(
seq_len: int,
image_starts: torch.Tensor,
num_image_tokens: int,
valid_lengths: torch.Tensor
) -> torch.Tensor

Boolean [batch, 1, seq, seq] mask: causal text, bidirectional image span, padded keys masked out.

Parameters:

seq_len
int

Padded sequence length.

image_starts
torch.Tensor

Long tensor of shape [batch], index of the first image token of every sample.

num_image_tokens
int

Number of image tokens (same for every sample of a batch).

valid_lengths
torch.Tensor

Long tensor of shape [batch], number of real (non-padding) tokens of every sample.

Returns: torch.Tensor

Boolean tensor of shape [batch, 1, sequence, sequence], indexed [batch, 1, query, key], true where

nemo_automodel.components.models.hunyuan_image3.model.build_moe_config(
overrides: dict[str, typing.Any] | None = None

MoE settings of the release: softmax over all experts, top-k, renormalized, one shared expert.

nemo_automodel.components.models.hunyuan_image3.model.ModelClass = HunyuanImage3ForCausalMM