nemo_automodel.components.checkpoint.state_dict_adapter
nemo_automodel.components.checkpoint.state_dict_adapter
Module Contents
Classes
API
One part of a checkpoint that is loaded and finished before the next part.
DCP fills checkpoint_tensors using their Hugging Face names and layouts. finish then converts any
temporary checkpoint tensors into the final model tensors named by model_keys before the loader advances to
the next part. Tensors that already have the model’s exact dtype and layout may point directly at model storage.
Abstract base class for state dict transformations.
This class defines the interface for converting between native model state dict format and other model state dict formats.
Most custom models need an adapter. Models whose HF weight names and tensor layouts already match can omit it.
Whether DCP can load the checkpoint with zero or small temporary tensors.
Most checkpoint tensors must load directly into the model’s existing weight memory. Small temporary tensors are allowed when they are converted and released after the read. Enable this only when the extra memory is safely below a full model copy, including when loading on one GPU.
Convert a single tensor from native format to HuggingFace format.
Parameters:
Fully qualified name of the tensor in native format
The tensor to convert
Additional arguments for conversion
Returns: list[tuple[str, Any]]
List of (fqn, tensor) tuples in HuggingFace format.
Obtain native model state dict from HuggingFace format.
Parameters:
The HuggingFace format state dict
Optional device mesh for DTensor expert parallelism. If provided, only loads experts needed for the current rank.
Returns: dict[str, Any]
The converted native model state dict
Return the Hugging Face keys produced by to_hf without converting real weights.
Parameters:
Native model state mapping. Tensor values may have arbitrary rank and axis order and retain their exact parameter or buffer layouts.
Returns: list[str]
Hugging Face state-dict keys in adapter iteration order.
Optionally load and finish a quantized checkpoint in small parts.
Ordinary adapters do not need to implement this method. It is only for adapters that cannot let DCP write checkpoint tensors directly into model storage because a dtype or layout conversion is required, but can complete that conversion for one small part of the model at a time.
Parameters:
Native model names mapped to the final parameter and persistent-buffer tensors that the checkpoint must populate. Tensors retain their model-specific shapes, axis orders, dtypes, devices, strides, distributed placements, and storage.
Optional device mesh describing distributed model storage.
Returns: Iterator[CheckpointLoadPart] | None
An iterator of dependency-complete load parts, or None when the adapter does not support this path.
Translate a PEFT target-module name to the HuggingFace layout.
adapter_config.json’s target_modules are collected from native module
names. Adapters whose to_hf renames modules (e.g. Kimi K3’s
mlp.experts.{E}.gate_proj -> block_sparse_moe.experts.{E}.w1)
should override this with the same renames so PEFT can resolve the
entries against the converted checkpoint.
Parameters:
A target-module name in native layout.
Whether to target the legacy Transformers v4 layout.
Returns: str
The name in HuggingFace layout. Defaults to the name unchanged.
Convert from native model state dict to HuggingFace format.
Parameters:
The native model state dict
Returns: dict[str, Any]
The converted HuggingFace format state dict