nemo_automodel.components.models.hunyuan_image3

View as Markdown

Submodules

Package Contents

Classes

NameDescription
HunyuanImage3ConfigConfiguration of the HunyuanImage-3.0 multimodal MoE transformer.
HunyuanImage3ForCausalMMHunyuanImage-3.0 transformer with the text-to-image flow-matching forward.

API

class nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config(
vocab_size: int = 133120,
hidden_size: int = 4096,
intermediate_size: int = 3072,
moe_intermediate_size: int | list[int] = 3072,
num_hidden_layers: int = 32,
num_attention_heads: int = 32,
num_key_value_heads: int = 8,
attention_head_dim: int = 128,
num_experts: int | list[int] = 64,
moe_topk: int | list[int] = 8,
num_shared_expert: int | list[int] = 1,
use_mixed_mlp_moe: bool = True,
norm_topk_prob: bool = True,
router_aux_loss_coef: float = 0.0,
hidden_act: str = 'silu',
rms_norm_eps: float = 1e-05,
rope_theta: float = 10000.0,
use_qk_norm: bool = True,
attention_bias: bool = False,
mlp_bias: bool = False,
max_position_embeddings: int = 22800,
img_proj_type: str = 'unet',
patch_size: int = 1,
patch_embed_hidden_dim: int = 1024,
image_base_size: int = 1024,
image_token_id: int = 128006,
timestep_token_id: int = 128017,
vae: dict[str, typing.Any] | None = None,
pad_token_id: int | None = 128009,
bos_token_id: int | None = 127958,
eos_token_id: int | list[int] | None = 127957,
tie_word_embeddings: bool = False,
torch_dtype: str = 'bfloat16',
kwargs: typing.Any = {}
)

Bases: PretrainedConfig

Configuration of the HunyuanImage-3.0 multimodal MoE transformer.

Architecture (tencent/HunyuanImage-3.0):

  • 32 decoder layers, hidden size 4096, GQA with 32 query / 8 KV heads, head dim 128
  • Every layer is MoE: 64 routed experts (top-8, softmax, renormalized) plus one shared expert
  • Per-head QK RMSNorm applied after RoPE; 2D RoPE over the joint text / image sequence
  • Image tokens come from VAE latents through a UNet patch embedding, and the diffusion velocity is read out through a UNet final layer, both conditioned on the flow-matching timestep
head_dim
int
keys_to_ignore_at_inference
= ['past_key_values']
latent_channels
int
model_type
= 'hunyuan_image_3_moe'
vae
class nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM(
kwargs: typing.Any = {}
)

Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin

HunyuanImage-3.0 transformer with the text-to-image flow-matching forward.

forward predicts the flow velocity of noisy latents given the token sequence; forward_text returns next-token logits (the release’s gen_text mode) and is used for parity checks.

_keep_in_fp32_modules_strict
= ['mlp.gate']
backend
= backend or BackendConfig()
final_layer
lm_head
model
patch_embed
state_dict_adapter
tie_word_embeddings_support
TieSupport = TieSupport.UNTIED_ONLY
time_embed
= TimestepEmbedder(hidden, dtype=dtype)
time_embed_2
= TimestepEmbedder(hidden, dtype=dtype)
timestep_emb
= TimestepEmbedder(hidden, dtype=dtype)
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward(
input_ids: torch.Tensor,
latents: torch.Tensor,
timestep: torch.Tensor,
valid_lengths: torch.Tensor | None = None,
return_dict: bool = False,
kwargs: typing.Any = {}
) -> tuple[torch.Tensor] | dict[str, torch.Tensor]

Predict the flow velocity of the noisy latents.

Parameters:

input_ids
torch.Tensor

[batch, seq] token ids laid out as described in the module docstring, right-padded. Every row holds exactly h*w contiguous image tokens, preceded by the <timestep> token.

latents
torch.Tensor

[batch, channels, h, w] noisy VAE latents.

timestep
torch.Tensor

[batch] flow-matching timesteps in [0, 1000] (sigma * 1000).

valid_lengths
torch.Tensor | NoneDefaults to None

[batch] number of non-padding tokens per row; defaults to the full length.

return_dict
boolDefaults to False

Return {"sample": velocity} instead of (velocity,).

Returns: tuple[torch.Tensor] | dict[str, torch.Tensor]

[batch, channels, h, w] predicted velocity, as a one-element tuple or a dict.

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.forward_text(
input_ids: torch.Tensor
) -> torch.Tensor

Next-token logits of a causal text-only sequence (the release’s gen_text mode).

Parameters:

input_ids
torch.Tensor

Long tensor of shape [batch, sequence].

Returns: torch.Tensor

fp32 tensor of shape [batch, sequence, vocab].

nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_config(
kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
classmethod
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.from_pretrained(
pretrained_model_name_or_path: str,
model_args: typing.Any = (),
kwargs: typing.Any = {}
) -> 'HunyuanImage3ForCausalMM'
classmethod
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_input_embeddings() -> torch.nn.Embedding
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.get_output_embeddings() -> torch.nn.Module
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.initialize_weights(
buffer_device: torch.device | None = None,
dtype: torch.dtype = torch.bfloat16
) -> None
nemo_automodel.components.models.hunyuan_image3.model.HunyuanImage3ForCausalMM.update_moe_gate_bias() -> None