nemo_automodel.components.models.hunyuan_image3.state_dict_adapter
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter
State dict conversion between the tencent/HunyuanImage-3.0 checkpoint and Automodel’s native layout.
Released checkpoint (HF) Native model.wte.weight model.embed_tokens.weight model.ln_f.weight model.norm.weight model.layers.{L}.mlp.gate.wg.weight [E, H] model.layers.{L}.mlp.gate.weight model.layers.{L}.mlp.shared_mlp.* model.layers.{L}.shared_mlp.* model.layers.{L}.mlp.experts.{e}.gate_and_up_proj.weight model.layers.{L}.mlp.experts.gate_and_up_projs [2I, H], rows = [up; gate] [E, H, 2I], columns = [gate | up] model.layers.{L}.mlp.experts.{e}.down_proj.weight [H, I] model.layers.{L}.mlp.experts.down_projs [E, I, H]
The routed experts go through the shared per-expert split / merge of MoESplitExpertsStateDictMixin, which works
with separate gate_proj / up_proj keys; this adapter fuses them into the released [up; gate] tensor on
the way out and splits them on the way in. On a checkpoint load the mixin hands out views into the model weight,
but DCP cannot write one fused checkpoint tensor through two views, so each fused tensor gets a host buffer whose
halves from_hf copies into the views; the grouped tensor then counts as loaded in place.
The VAE (vae.*) and the vision encoder (vision_model.*, vision_aligner.*) of the release are not part of
the training model; their keys are dropped on load and absent on save.
Module Contents
Classes
Functions
Data
API
Bases: MoESplitExpertsStateDictMixin, StateDictAdapter
Converts between the released HunyuanImage-3.0 checkpoint and the native grouped-expert model.
Convert one native tensor to one or more released checkpoint entries.
Parameters:
Native parameter name.
Native tensor; grouped gate_and_up_projs have shape [experts, hidden, 2 * expert_hidden]
with columns [gate | up] and down_projs [experts, expert_hidden, hidden] (DTensors sharded
on the expert axis under expert parallelism).
Forwarded to the shared expert split (exclude_key_regex filters the output keys).
Returns: list[tuple[str, Any]]
(key, tensor) entries in the released layout: per local expert, gate_and_up_proj of shape
Convert released checkpoint entries to the native layout (local experts only under EP).
Parameters:
Released entries; consumed (popped) by this call. Per-expert gate_and_up_proj has
shape [2 * expert_hidden, hidden] with rows [up; gate].
Mesh whose ep axis selects the local experts, or None for all experts.
Unused; accepted for the base-class signature.
Returns: dict[str, Any]
Native state dict; grouped gate_and_up_projs of shape [local_experts, hidden, 2 * expert_hidden]
Match PEFT target names to the released fused projections.
Parameters:
Target-module path after the shared exporter expands combined projections.
Legacy export selection; both formats use the same released module names.
Returns: str
Released module path, with shared experts renamed and split QKV targets reunited.
Convert a native state dict to released checkpoint keys.
With for_checkpoint_load=True the fused expert entries become host buffers for DCP (one
[2 * expert_hidden, hidden] tensor per local expert, 50 MB in bf16 for the release, about 13 GB per rank with
8 local experts) that from_hf copies into the model weight; a new load conversion forgets the views
of an earlier one.
Check whether plain checkpoint destinations alias the grouped model weight.
Parameters:
Per-expert entries with tensors of shape [expert_hidden, hidden]. DTensors retain a remaining mesh dimension and must use the distributed conversion path, even when their local storage aliases.
Grouped tensor of shape [experts, hidden, 2 * expert_hidden], possibly a DTensor.
Returns: bool
Whether every destination is a plain tensor aliasing the source’s local storage.
Pair per-expert up_proj / gate_proj entries by expert.
Parameters:
(key, tensor) entries; the gate/up tensors have shape [expert_hidden, hidden].
Returns: tuple[dict[str, tuple[Any, Any]], list[tuple[str, Any]]]
({expert stem: (up, gate)}, other entries).