ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsHunyuan Image3nemo_automodel.components.models.hunyuan_image3.state_dict_adapter

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter

View as Markdown

State dict conversion between the tencent/HunyuanImage-3.0 checkpoint and Automodel’s native layout.

Released checkpoint (HF) Native model.wte.weight model.embed_tokens.weight model.ln_f.weight model.norm.weight model.layers.{L}.mlp.gate.wg.weight [E, H] model.layers.{L}.mlp.gate.weight model.layers.{L}.mlp.shared_mlp.* model.layers.{L}.shared_mlp.* model.layers.{L}.mlp.experts.{e}.gate_and_up_proj.weight model.layers.{L}.mlp.experts.gate_and_up_projs [2I, H], rows = [up; gate] [E, H, 2I], columns = [gate | up] model.layers.{L}.mlp.experts.{e}.down_proj.weight [H, I] model.layers.{L}.mlp.experts.down_projs [E, I, H]

The routed experts go through the shared per-expert split / merge of MoESplitExpertsStateDictMixin, which works with separate gate_proj / up_proj keys; this adapter fuses them into the released [up; gate] tensor on the way out and splits them on the way in. On a checkpoint load the mixin hands out views into the model weight, but DCP cannot write one fused checkpoint tensor through two views, so each fused tensor gets a host buffer whose halves from_hf copies into the views; the grouped tensor then counts as loaded in place.

The VAE (vae.*) and the vision encoder (vision_model.*, vision_aligner.*) of the release are not part of the training model; their keys are dropped on load and absent on save.

Module Contents

Classes

NameDescription
HunyuanImage3StateDictAdapterConverts between the released HunyuanImage-3.0 checkpoint and the native grouped-expert model.

Functions

NameDescription
_all_aliasCheck whether plain checkpoint destinations alias the grouped model weight.
_group_split_expertsPair per-expert up_proj / gate_proj entries by expert.
_rename-
_rename_table-

Data

_FUSED_EXPERT_KEY

_HF_TO_NATIVE_RENAMES

_NATIVE_TO_HF_RENAMES

_RENAMES

_SPLIT_EXPERT_KEY

_UNUSED_HF_PREFIXES

API

class nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter(
config: typing.Any,
dtype: torch.dtype = torch.bfloat16
)

Bases: MoESplitExpertsStateDictMixin, StateDictAdapter

Converts between the released HunyuanImage-3.0 checkpoint and the native grouped-expert model.

_fused_load_destinations
dict[str, tuple[Tensor, Tensor]] = {}
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.convert_single_tensor_to_hf(
fqn: str,
tensor: typing.Any,
kwargs: typing.Any = {}
) -> list[tuple[str, typing.Any]]

Convert one native tensor to one or more released checkpoint entries.

Parameters:

fqn
str

Native parameter name.

tensor
Any

Native tensor; grouped gate_and_up_projs have shape [experts, hidden, 2 * expert_hidden] with columns [gate | up] and down_projs [experts, expert_hidden, hidden] (DTensors sharded on the expert axis under expert parallelism).

**kwargs
AnyDefaults to {}

Forwarded to the shared expert split (exclude_key_regex filters the output keys).

Returns: list[tuple[str, Any]]

(key, tensor) entries in the released layout: per local expert, gate_and_up_proj of shape

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.from_hf(
hf_state_dict: dict[str, typing.Any],
device_mesh: torch.distributed.device_mesh.DeviceMesh | None = None,
kwargs: typing.Any = {}
) -> dict[str, typing.Any]

Convert released checkpoint entries to the native layout (local experts only under EP).

Parameters:

hf_state_dict
dict[str, Any]

Released entries; consumed (popped) by this call. Per-expert gate_and_up_proj has shape [2 * expert_hidden, hidden] with rows [up; gate].

device_mesh
DeviceMesh | NoneDefaults to None

Mesh whose ep axis selects the local experts, or None for all experts.

**kwargs
AnyDefaults to {}

Unused; accepted for the base-class signature.

Returns: dict[str, Any]

Native state dict; grouped gate_and_up_projs of shape [local_experts, hidden, 2 * expert_hidden]

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.map_peft_target_module_to_hf(
name: str,
v4_compatible: bool = False
) -> str

Match PEFT target names to the released fused projections.

Parameters:

name
str

Target-module path after the shared exporter expands combined projections.

v4_compatible
boolDefaults to False

Legacy export selection; both formats use the same released module names.

Returns: str

Released module path, with shared experts renamed and split QKV targets reunited.

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.to_hf(
state_dict: dict[str, typing.Any],
exclude_key_regex: str | None = None,
kwargs: typing.Any = {}
) -> dict[str, typing.Any]

Convert a native state dict to released checkpoint keys.

With for_checkpoint_load=True the fused expert entries become host buffers for DCP (one [2 * expert_hidden, hidden] tensor per local expert, 50 MB in bf16 for the release, about 13 GB per rank with 8 local experts) that from_hf copies into the model weight; a new load conversion forgets the views of an earlier one.

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._all_alias(
pairs: list[tuple[str, typing.Any]],
tensor: typing.Any
) -> bool

Check whether plain checkpoint destinations alias the grouped model weight.

Parameters:

pairs
list[tuple[str, Any]]

Per-expert entries with tensors of shape [expert_hidden, hidden]. DTensors retain a remaining mesh dimension and must use the distributed conversion path, even when their local storage aliases.

tensor
Any

Grouped tensor of shape [experts, hidden, 2 * expert_hidden], possibly a DTensor.

Returns: bool

Whether every destination is a plain tensor aliasing the source’s local storage.

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._group_split_experts(
pairs: list[tuple[str, typing.Any]]
) -> tuple[dict[str, tuple[typing.Any, typing.Any]], list[tuple[str, typing.Any]]]

Pair per-expert up_proj / gate_proj entries by expert.

Parameters:

pairs
list[tuple[str, Any]]

(key, tensor) entries; the gate/up tensors have shape [expert_hidden, hidden].

Returns: tuple[dict[str, tuple[Any, Any]], list[tuple[str, Any]]]

({expert stem: (up, gate)}, other entries).

nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._rename(
key: str,
renames: tuple[tuple[re.Pattern[str], str], ...]
) -> str
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._rename_table(
src: int,
dst: int
) -> tuple[tuple[re.Pattern[str], str], ...]
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._FUSED_EXPERT_KEY = re.compile('^(?P<stem>.*\\.mlp\\.experts\\.\\d+)\\.gate_and_up_proj\\.weight$')
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._HF_TO_NATIVE_RENAMES = _rename_table(1, 0)
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._NATIVE_TO_HF_RENAMES = _rename_table(0, 1)
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._RENAMES: tuple[tuple[str, str], ...] = (('^model\\.embed_tokens\\.weight$', '^model\\.wte\\.weight$'), ('^model\\.norm\...
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._SPLIT_EXPERT_KEY = re.compile('^(?P<stem>.*\\.mlp\\.experts\\.\\d+)\\.(?P<proj>gate_proj|up_proj)\\...
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._UNUSED_HF_PREFIXES = ('vae.', 'vision_model.', 'vision_aligner.')