nemo_automodel.components.models.hunyuan_image3.config
nemo_automodel.components.models.hunyuan_image3.config
Configuration for tencent/HunyuanImage-3.0.
The checkpoint’s config.json carries auto_map entries that point at remote code. Registering this class
for model_type="hunyuan_image_3_moe" lets AutoConfig read the checkpoint without trust_remote_code.
Fields not listed here (VAE / ViT settings, special token ids, …) are kept as plain attributes by
PretrainedConfig.
Module Contents
Classes
Functions
API
Bases: PretrainedConfig
Configuration of the HunyuanImage-3.0 multimodal MoE transformer.
Architecture (tencent/HunyuanImage-3.0):
- 32 decoder layers, hidden size 4096, GQA with 32 query / 8 KV heads, head dim 128
- Every layer is MoE: 64 routed experts (top-8, softmax, renormalized) plus one shared expert
- Per-head QK RMSNorm applied after RoPE; 2D RoPE over the joint text / image sequence
- Image tokens come from VAE latents through a UNet patch embedding, and the diffusion velocity is read out through a UNet final layer, both conditioned on the flow-matching timestep
head_dim
keys_to_ignore_at_inference
latent_channels
model_type
vae
Resolve a config field that is either a scalar or a per-layer list.