nemo_automodel.components.models.hunyuan_image3
nemo_automodel.components.models.hunyuan_image3
Submodules
nemo_automodel.components.models.hunyuan_image3.confignemo_automodel.components.models.hunyuan_image3.flow_adapternemo_automodel.components.models.hunyuan_image3.layersnemo_automodel.components.models.hunyuan_image3.modelnemo_automodel.components.models.hunyuan_image3.pipelinenemo_automodel.components.models.hunyuan_image3.releasenemo_automodel.components.models.hunyuan_image3.ropenemo_automodel.components.models.hunyuan_image3.state_dict_adapter
Package Contents
Classes
API
Bases: PretrainedConfig
Configuration of the HunyuanImage-3.0 multimodal MoE transformer.
Architecture (tencent/HunyuanImage-3.0):
- 32 decoder layers, hidden size 4096, GQA with 32 query / 8 KV heads, head dim 128
- Every layer is MoE: 64 routed experts (top-8, softmax, renormalized) plus one shared expert
- Per-head QK RMSNorm applied after RoPE; 2D RoPE over the joint text / image sequence
- Image tokens come from VAE latents through a UNet patch embedding, and the diffusion velocity is read out through a UNet final layer, both conditioned on the flow-matching timestep
Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin
HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.
forward predicts the flow velocity of noisy latents given the token sequence; forward_text returns
next-token logits (the release’s gen_text mode) and is used for parity checks.
Predict the flow velocity of the noisy latents.
Parameters:
[batch, seq] token ids laid out as described in the module docstring, right-padded.
Every row holds exactly h*w contiguous image tokens, preceded by the <timestep> token.
[batch, channels, h, w] noisy VAE latents.
[batch] flow-matching timesteps in [0, 1000] (sigma * 1000).
[batch] number of non-padding tokens per row; defaults to the full length.
Return {"sample": velocity} instead of (velocity,).
Returns: tuple[torch.Tensor] | dict[str, torch.Tensor]
[batch, channels, h, w] predicted velocity, as a one-element tuple or a dict.
Next-token logits of a causal text-only sequence (the release’s gen_text mode).
Parameters:
Long tensor of shape [batch, sequence].
Returns: torch.Tensor
fp32 tensor of shape [batch, sequence, vocab].