nemo_automodel.components.models.hunyuan_image3.config

View as Markdown

Configuration for tencent/HunyuanImage-3.0.

The checkpoint’s config.json carries auto_map entries that point at remote code. Registering this class for model_type="hunyuan_image_3_moe" lets AutoConfig read the checkpoint without trust_remote_code. Fields not listed here (VAE / ViT settings, special token ids, …) are kept as plain attributes by PretrainedConfig.

Module Contents

Classes

NameDescription
HunyuanImage3ConfigConfiguration of the HunyuanImage-3.0 multimodal MoE transformer.

Functions

NameDescription
per_layerResolve a config field that is either a scalar or a per-layer list.

API

class nemo_automodel.components.models.hunyuan_image3.config.HunyuanImage3Config(
vocab_size: int = 133120,
hidden_size: int = 4096,
intermediate_size: int = 3072,
moe_intermediate_size: int | list[int] = 3072,
num_hidden_layers: int = 32,
num_attention_heads: int = 32,
num_key_value_heads: int = 8,
attention_head_dim: int = 128,
num_experts: int | list[int] = 64,
moe_topk: int | list[int] = 8,
num_shared_expert: int | list[int] = 1,
use_mixed_mlp_moe: bool = True,
norm_topk_prob: bool = True,
router_aux_loss_coef: float = 0.0,
hidden_act: str = 'silu',
rms_norm_eps: float = 1e-05,
rope_theta: float = 10000.0,
use_qk_norm: bool = True,
attention_bias: bool = False,
mlp_bias: bool = False,
max_position_embeddings: int = 22800,
img_proj_type: str = 'unet',
patch_size: int = 1,
patch_embed_hidden_dim: int = 1024,
image_base_size: int = 1024,
image_token_id: int = 128006,
timestep_token_id: int = 128017,
vae: dict[str, typing.Any] | None = None,
pad_token_id: int | None = 128009,
bos_token_id: int | None = 127958,
eos_token_id: int | list[int] | None = 127957,
tie_word_embeddings: bool = False,
torch_dtype: str = 'bfloat16',
kwargs: typing.Any = {}
)

Bases: PretrainedConfig

Configuration of the HunyuanImage-3.0 multimodal MoE transformer.

Architecture (tencent/HunyuanImage-3.0):

  • 32 decoder layers, hidden size 4096, GQA with 32 query / 8 KV heads, head dim 128
  • Every layer is MoE: 64 routed experts (top-8, softmax, renormalized) plus one shared expert
  • Per-head QK RMSNorm applied after RoPE; 2D RoPE over the joint text / image sequence
  • Image tokens come from VAE latents through a UNet patch embedding, and the diffusion velocity is read out through a UNet final layer, both conditioned on the flow-matching timestep
head_dim
int
keys_to_ignore_at_inference
= ['past_key_values']
latent_channels
int
model_type
= 'hunyuan_image_3_moe'
vae
nemo_automodel.components.models.hunyuan_image3.config.per_layer(
value: int | list[int],
layer_idx: int
) -> int

Resolve a config field that is either a scalar or a per-layer list.