> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.hunyuan_image3.state_dict_adapter

State dict conversion between the tencent/HunyuanImage-3.0 checkpoint and Automodel's native layout.

Released checkpoint (HF)                                   Native
model.wte.weight                                           model.embed\_tokens.weight
model.ln\_f.weight                                          model.norm.weight
model.layers.\{L}.mlp.gate.wg.weight          \[E, H]        model.layers.\{L}.mlp.gate.weight
model.layers.\{L}.mlp.shared\_mlp.\*                          model.layers.\{L}.shared\_mlp.\*
model.layers.\{L}.mlp.experts.\{e}.gate\_and\_up\_proj.weight   model.layers.\{L}.mlp.experts.gate\_and\_up\_projs
\[2I, H], rows = \[up; gate]                                 \[E, H, 2I], columns = \[gate | up]
model.layers.\{L}.mlp.experts.\{e}.down\_proj.weight \[H, I]   model.layers.\{L}.mlp.experts.down\_projs \[E, I, H]

The routed experts go through the shared per-expert split / merge of `MoESplitExpertsStateDictMixin`, which works
with separate `gate_proj` / `up_proj` keys; this adapter fuses them into the released `[up; gate]` tensor on
the way out and splits them on the way in. On a checkpoint load the mixin hands out views into the model weight,
but DCP cannot write one fused checkpoint tensor through two views, so each fused tensor gets a host buffer whose
halves `from_hf` copies into the views; the grouped tensor then counts as loaded in place.

The VAE (`vae.*`) and the vision encoder (`vision_model.*`, `vision_aligner.*`) of the release are not part of
the training model; their keys are dropped on load and absent on save.

## Module Contents

### Classes

| Name                                                                                                                                 | Description                                                                                    |
| ------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------- |
| [`HunyuanImage3StateDictAdapter`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-HunyuanImage3StateDictAdapter) | Converts between the released HunyuanImage-3.0 checkpoint and the native grouped-expert model. |

### Functions

| Name                                                                                                               | Description                                                                 |
| ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------- |
| [`_all_alias`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_all_alias)                     | Check whether plain checkpoint destinations alias the grouped model weight. |
| [`_group_split_experts`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_group_split_experts) | Pair per-expert `up_proj` / `gate_proj` entries by expert.                  |
| [`_rename`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_rename)                           | -                                                                           |
| [`_rename_table`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_rename_table)               | -                                                                           |

### Data

[`_FUSED_EXPERT_KEY`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_FUSED_EXPERT_KEY)

[`_HF_TO_NATIVE_RENAMES`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_HF_TO_NATIVE_RENAMES)

[`_NATIVE_TO_HF_RENAMES`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_NATIVE_TO_HF_RENAMES)

[`_RENAMES`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_RENAMES)

[`_SPLIT_EXPERT_KEY`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_SPLIT_EXPERT_KEY)

[`_UNUSED_HF_PREFIXES`](#nemo_automodel-components-models-hunyuan_image3-state_dict_adapter-_UNUSED_HF_PREFIXES)

### API

```python
class nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter(
    config: typing.Any,
    moe_config: nemo_automodel.components.moe.config.MoEConfig,
    backend: nemo_automodel.components.models.common.BackendConfig,
    dtype: torch.dtype = torch.bfloat16
)
```

**Bases:** [MoESplitExpertsStateDictMixin](/nemo-automodel/nemo_automodel/components/moe/state_dict_mixin#nemo_automodel-components-moe-state_dict_mixin-MoESplitExpertsStateDictMixin), [StateDictAdapter](/nemo-automodel/nemo_automodel/components/checkpoint/state_dict_adapter#nemo_automodel-components-checkpoint-state_dict_adapter-StateDictAdapter)

Converts between the released HunyuanImage-3.0 checkpoint and the native grouped-expert model.

**`_fused_load_destinations`** `dict[str, tuple[Tensor, Tensor]] = {}`

---

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.convert_single_tensor_to_hf(
    fqn: str,
    tensor: typing.Any,
    kwargs: typing.Any = {}
) -> list[tuple[str, typing.Any]]
```

Convert one native tensor to one or more released checkpoint entries.

**Parameters:**

**`fqn`** `str`

Native parameter name.

---

**`tensor`** `Any`

Native tensor; grouped `gate_and_up_projs` have shape \[experts, hidden, 2 \* expert\_hidden]
with columns `[gate | up]` and `down_projs` \[experts, expert\_hidden, hidden] (DTensors sharded
on the expert axis under expert parallelism).

---

**`**kwargs`** `Any` — default: \{}

Forwarded to the shared expert split (`exclude_key_regex` filters the output keys).

---

**Returns:** `list[tuple[str, Any]]`

`(key, tensor)` entries in the released layout: per local expert, `gate_and_up_proj` of shape

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.from_hf(
    hf_state_dict: dict[str, typing.Any],
    device_mesh: torch.distributed.device_mesh.DeviceMesh | None = None,
    kwargs: typing.Any = {}
) -> dict[str, typing.Any]
```

Convert released checkpoint entries to the native layout (local experts only under EP).

**Parameters:**

**`hf_state_dict`** `dict[str, Any]`

Released entries; consumed (popped) by this call. Per-expert `gate_and_up_proj` has
shape \[2 \* expert\_hidden, hidden] with rows `[up; gate]`.

---

**`device_mesh`** `DeviceMesh | None` — default: None

Mesh whose `ep` axis selects the local experts, or `None` for all experts.

---

**`**kwargs`** `Any` — default: \{}

Unused; accepted for the base-class signature.

---

**Returns:** `dict[str, Any]`

Native state dict; grouped `gate_and_up_projs` of shape \[local\_experts, hidden, 2 \* expert\_hidden]

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.map_peft_target_module_to_hf(
    name: str,
    v4_compatible: bool = False
) -> str
```

Match PEFT target names to the released fused projections.

**Parameters:**

**`name`** `str`

Target-module path after the shared exporter expands combined projections.

---

**`v4_compatible`** `bool` — default: False

Legacy export selection; both formats use the same released module names.

---

**Returns:** `str`

Released module path, with shared experts renamed and split QKV targets reunited.

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter.HunyuanImage3StateDictAdapter.to_hf(
    state_dict: dict[str, typing.Any],
    exclude_key_regex: str | None = None,
    kwargs: typing.Any = {}
) -> dict[str, typing.Any]
```

Convert a native state dict to released checkpoint keys.

With `for_checkpoint_load=True` the fused expert entries become host buffers for DCP (one
\[2 \* expert\_hidden, hidden] tensor per local expert, 50 MB in bf16 for the release, about 13 GB per rank with
8 local experts) that `from_hf` copies into the model weight; a new load conversion forgets the views
of an earlier one.

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._all_alias(
    pairs: list[tuple[str, typing.Any]],
    tensor: typing.Any
) -> bool
```

Check whether plain checkpoint destinations alias the grouped model weight.

**Parameters:**

**`pairs`** `list[tuple[str, Any]]`

Per-expert entries with tensors of shape \[expert\_hidden, hidden]. DTensors retain a remaining
mesh dimension and must use the distributed conversion path, even when their local storage aliases.

---

**`tensor`** `Any`

Grouped tensor of shape \[experts, hidden, 2 \* expert\_hidden], possibly a DTensor.

---

**Returns:** `bool`

Whether every destination is a plain tensor aliasing the source's local storage.

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._group_split_experts(
    pairs: list[tuple[str, typing.Any]]
) -> tuple[dict[str, tuple[typing.Any, typing.Any]], list[tuple[str, typing.Any]]]
```

Pair per-expert `up_proj` / `gate_proj` entries by expert.

**Parameters:**

**`pairs`** `list[tuple[str, Any]]`

`(key, tensor)` entries; the gate/up tensors have shape \[expert\_hidden, hidden].

---

**Returns:** `tuple[dict[str, tuple[Any, Any]], list[tuple[str, Any]]]`

`(&#123;expert stem: (up, gate)&#125;, other entries)`.

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._rename(
    key: str,
    renames: tuple[tuple[re.Pattern[str], str], ...]
) -> str
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._rename_table(
    src: int,
    dst: int
) -> tuple[tuple[re.Pattern[str], str], ...]
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._FUSED_EXPERT_KEY = re.compile('^(?P<stem>.*\\.mlp\\.experts\\.\\d+)\\.gate_and_up_proj\\.weight$')
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._HF_TO_NATIVE_RENAMES = _rename_table(1, 0)
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._NATIVE_TO_HF_RENAMES = _rename_table(0, 1)
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._RENAMES: tuple[tuple[str, str], ...] = (('^model\\.embed_tokens\\.weight$', '^model\\.wte\\.weight$'), ('^model\\.norm\...
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._SPLIT_EXPERT_KEY = re.compile('^(?P<stem>.*\\.mlp\\.experts\\.\\d+)\\.(?P<proj>gate_proj|up_proj)\\...
```

```python
nemo_automodel.components.models.hunyuan_image3.state_dict_adapter._UNUSED_HF_PREFIXES = ('vae.', 'vision_model.', 'vision_aligner.')
```