ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsDistributednemo_automodel.components.distributed.model_parallelizer

nemo_automodel.components.distributed.model_parallelizer

View as Markdown

Resolve and execute model-owned parallelization sidecars.

Module Contents

Classes

NameDescription
ModelParallelizerSingle model-owned parallelization sidecar contract.

Functions

NameDescription
_apply_model_parallelizerExecute shared strategy dispatch for one model-owned parallelizer.
_parallelize_ddp-
_parallelize_fsdp2-
_parallelize_megatron_fsdp-
_parallelize_moe-
_parallelize_unsharded_fsdp2-
compile_parallelized_modelCompile FSDP2 layers after parallelization when requested.
get_model_parallelizerReturn the class-owned sidecar, or the shared default implementation.
parallelize_modelApply all requested parallelisms through the model-owned contract.

Data

_DEFAULT_PARALLELIZER

API

class nemo_automodel.components.distributed.parallelizer.ModelParallelizer()

Single model-owned parallelization sidecar contract.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer._apply(
model: torch.nn.Module,
device_mesh: torch.distributed.device_mesh.DeviceMesh,
mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None = None,
offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
sequence_parallel: bool = False,
activation_checkpointing: bool = False,
tp_shard_plan: typing.Union[typing.Dict[str, torch.distributed.tensor.parallel.ParallelStyle], str] | None = None,
dp_replicate_mesh_name: str = 'dp_replicate',
dp_shard_cp_mesh_name: str = 'dp_shard_cp',
tp_mesh_name: str = 'tp',
enable_async_tensor_parallel: bool = False,
enable_compile: bool = False,
enable_fsdp2_prefetch: bool = True,
fsdp2_backward_prefetch_depth: int = 2,
fsdp2_forward_prefetch_depth: int = 1,
reshard_after_forward: bool | None = None,
activation_checkpointing_scope: nemo_automodel.components.distributed.config.ActivationCheckpointingScope | None = 'all',
reapply_trainability: collections.abc.Callable[[nn.Module], None] | None = None
) -> torch.nn.Module

Apply the shared dense FSDP2 implementation.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer._apply_fsdp_sharding(
module: torch.nn.Module,
mesh: torch.distributed.device_mesh.DeviceMesh,
mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None,
offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
enable_fsdp2_prefetch: bool = True,
fsdp2_backward_prefetch_depth: int = 2,
fsdp2_forward_prefetch_depth: int = 1,
reshard_after_forward: bool | None = None,
ignored_multimodal_params: set[torch.nn.Parameter] | None = None
) -> None

Wrap the model’s submodules into FSDP2 units.

Model-owned sidecar strategies deriving from this class override this hook to change how parameters are grouped into FSDP units without reimplementing the surrounding TP/AC/mixed-precision flow. Override _fully_shard_module when a model needs a different primitive.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer._fully_shard_module(
module: torch.nn.Module,
kwargs = {}
) -> torch.nn.Module

Apply the FSDP2 primitive used by this model sidecar.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer._use_full_layer_activation_checkpointing(
model: torch.nn.Module
) -> bool

Return whether this model safely opts into whole-layer checkpointing.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer._validate_tp_mesh(
model: torch.nn.Module,
tp_mesh: torch.distributed.device_mesh.DeviceMesh
) -> None

Validate the model’s attention topology against its TP mesh.

nemo_automodel.components.distributed.parallelizer.ModelParallelizer.parallelize(
model: torch.nn.Module,
) -> torch.nn.Module

Apply every requested parallelism and return the parallelized model.

nemo_automodel.components.distributed.model_parallelizer._apply_model_parallelizer(
model: torch.nn.Module,
) -> torch.nn.Module

Execute shared strategy dispatch for one model-owned parallelizer.

nemo_automodel.components.distributed.model_parallelizer._parallelize_ddp(
model: torch.nn.Module,
) -> torch.nn.Module
nemo_automodel.components.distributed.model_parallelizer._parallelize_fsdp2(
model: torch.nn.Module,
) -> torch.nn.Module
nemo_automodel.components.distributed.model_parallelizer._parallelize_megatron_fsdp(
model: torch.nn.Module,
) -> torch.nn.Module
nemo_automodel.components.distributed.model_parallelizer._parallelize_moe(
model: torch.nn.Module,
) -> torch.nn.Module
nemo_automodel.components.distributed.model_parallelizer._parallelize_unsharded_fsdp2(
model: torch.nn.Module,
) -> torch.nn.Module
nemo_automodel.components.distributed.model_parallelizer.compile_parallelized_model(
model: torch.nn.Module,
) -> None

Compile FSDP2 layers after parallelization when requested.

nemo_automodel.components.distributed.model_parallelizer.get_model_parallelizer(
model: torch.nn.Module

Return the class-owned sidecar, or the shared default implementation.

nemo_automodel.components.distributed.model_parallelizer.parallelize_model(
model: torch.nn.Module,
) -> torch.nn.Module

Apply all requested parallelisms through the model-owned contract.

nemo_automodel.components.distributed.model_parallelizer._DEFAULT_PARALLELIZER = ModelParallelizer()