nemo_automodel.components.models.hunyuan_image3.model
nemo_automodel.components.models.hunyuan_image3.model
HunyuanImage-3.0 (tencent/HunyuanImage-3.0) for flow-matching text-to-image training.
The released model is a single 80B-total / 13B-active MoE decoder that handles text and images in one sequence. For text-to-image generation the sequence is::
<bos> prompt <boi> <img_size_> <img_ratio_> <timestep> <img> * (h*w) <eoi>
The <timestep> slot receives a timestep embedding and the <img> slots receive the noisy VAE latents through
a UNet patch embedding. Text tokens attend causally; the image tokens attend to each other bidirectionally. The
final hidden states at the image slots (before the final norm) go through a UNet final layer that predicts the
flow velocity noise - x0.
The decoder uses Automodel’s MoE under model.layers[*].mlp so
the MoE parallelizer can apply FSDP2 and expert parallelism. The VAE and the vision encoder of the release are not
part of this module: latents are precomputed during preprocessing, and image conditioning is not supported.
Module Contents
Classes
Functions
Data
API
Bases: Module
Pre-norm decoder layer: GQA attention, then routed experts (MoE) plus the release’s fused shared expert.
The shared expert lives in the block, not in MoE, because its fused [up; gate] projection is not the
MLP that MoE builds and manages.
Run one decoder layer.
Parameters:
Tensor of shape [batch, sequence, hidden].
fp32 tensor of shape [batch, sequence, head_dim] rotary table.
fp32 tensor of shape [batch, sequence, head_dim] rotary table.
Boolean tensor of shape [batch, 1, sequence, sequence], true where attention is
allowed, or None for causal attention.
Boolean tensor of shape [batch, sequence], true at padding, or None.
Returns: torch.Tensor
Tensor of shape [batch, sequence, hidden].
Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin
HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.
forward predicts the flow velocity of noisy latents given the token sequence; forward_text returns
next-token logits (the release’s gen_text mode) and is used for parity checks.
Predict the flow velocity of the noisy latents.
Parameters:
[batch, seq] token ids laid out as described in the module docstring, right-padded.
Every row holds exactly h*w contiguous image tokens, preceded by the <timestep> token.
[batch, channels, h, w] noisy VAE latents.
[batch] flow-matching timesteps in [0, 1000] (sigma * 1000).
[batch] number of non-padding tokens per row; defaults to the full length.
Return {"sample": velocity} instead of (velocity,).
Returns: tuple[torch.Tensor] | dict[str, torch.Tensor]
[batch, channels, h, w] predicted velocity, as a one-element tuple or a dict.
Next-token logits of a causal text-only sequence (the release’s gen_text mode).
Parameters:
Long tensor of shape [batch, sequence].
Returns: torch.Tensor
fp32 tensor of shape [batch, sequence, vocab].
Bases: Module
Decoder backbone: token embedding, MoE layers and the final norm (wte / ln_f in the release).
Run the decoder layers.
Parameters:
Tensor of shape [batch, sequence, hidden].
fp32 tensor of shape [batch, sequence, head_dim] rotary table.
fp32 tensor of shape [batch, sequence, head_dim] rotary table.
Boolean tensor of shape [batch, 1, sequence, sequence], true where attention is
allowed, or None for causal attention.
Boolean tensor of shape [batch, sequence], true at padding, or None.
Returns: torch.Tensor
Tensor of shape [batch, sequence, hidden]: the last hidden states before the final norm.
Parameter dtype from the config (dtype under transformers 5, torch_dtype before).
Boolean [batch, 1, seq, seq] mask: causal text, bidirectional image span, padded keys masked out.
Parameters:
Padded sequence length.
Long tensor of shape [batch], index of the first image token of every sample.
Number of image tokens (same for every sample of a batch).
Long tensor of shape [batch], number of real (non-padding) tokens of every sample.
Returns: torch.Tensor
Boolean tensor of shape [batch, 1, sequence, sequence], indexed [batch, 1, query, key], true where
MoE settings of the release: softmax over all experts, top-k, renormalized, one shared expert.