nemo_automodel.components.checkpoint.state_dict_adapter

View as Markdown

Module Contents

Classes

NameDescription
StateDictAdapterAbstract base class for state dict transformations.

API

class nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter()
Abstract

Abstract base class for state dict transformations.

This class defines the interface for converting between native model state dict format and other model state dict formats.

_supports_checkpoint_load_without_full_copy
bool = False
_supports_write_through_checkpoint_load
bool = False
supports_checkpoint_load_without_full_copy
bool

Whether DCP can load this adapter without another full set of model weights.

Large checkpoint tensors must be loaded into the model’s existing weight memory. Small temporary tensors are allowed when they can be applied and discarded without making a model-sized copy. For example, Gemma4 loads a scale tensor and applies it to already-loaded expert weights.

supports_write_through_checkpoint_load
bool

Whether every checkpoint tensor is loaded directly into the model’s existing weight memory.

Enable this only when writing every tensor returned by to_hf for base-checkpoint loading updates the model itself. This lets the loader skip a complete CPU copy of the checkpoint.

nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter.convert_single_tensor_to_hf(
fqn: str,
tensor: typing.Any,
kwargs = {}
) -> list[tuple[str, typing.Any]]
abstract

Convert a single tensor from native format to HuggingFace format.

Parameters:

fqn
str

Fully qualified name of the tensor in native format

tensor
Any

The tensor to convert

**kwargs
Defaults to {}

Additional arguments for conversion

Returns: list[tuple[str, Any]]

List of (fqn, tensor) tuples in HuggingFace format.

nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter.from_hf(
hf_state_dict: dict[str, typing.Any],
device_mesh: typing.Optional[torch.distributed.device_mesh.DeviceMesh] = None,
kwargs = {}
) -> dict[str, typing.Any]
abstract

Obtain native model state dict from HuggingFace format.

Parameters:

hf_state_dict
dict[str, Any]

The HuggingFace format state dict

device_mesh
Optional[DeviceMesh]Defaults to None

Optional device mesh for DTensor expert parallelism. If provided, only loads experts needed for the current rank.

Returns: dict[str, Any]

The converted native model state dict

nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter.get_hf_state_dict_keys(
state_dict: dict[str, typing.Any]
) -> list[str]

Return the Hugging Face keys produced by to_hf.

Parameters:

state_dict
dict[str, Any]

Native model state mapping. Tensor values may have arbitrary rank and axis order and retain their exact parameter or buffer layouts.

Returns: list[str]

Hugging Face state-dict keys in adapter iteration order.

nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter.map_peft_target_module_to_hf(
name: str
) -> str

Translate a PEFT target-module name to the HuggingFace layout.

adapter_config.json’s target_modules are collected from native module names. Adapters whose to_hf renames modules (e.g. Kimi K3’s mlp.experts.{E}.gate_proj -> block_sparse_moe.experts.{E}.w1) should override this with the same renames so PEFT can resolve the entries against the converted checkpoint.

Parameters:

name
str

A target-module name in native layout.

Returns: str

The name in HuggingFace layout. Defaults to the name unchanged.

nemo_automodel.components.checkpoint.state_dict_adapter.StateDictAdapter.to_hf(
state_dict: dict[str, typing.Any],
kwargs = {}
) -> dict[str, typing.Any]
abstract

Convert from native model state dict to HuggingFace format.

Parameters:

state_dict
dict[str, Any]

The native model state dict

Returns: dict[str, Any]

The converted HuggingFace format state dict