> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.distributed.model_parallelizer

Resolve and execute model-owned parallelization sidecars.

## Module Contents

### Classes

| Name                                                                                         | Description                                          |
| -------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
| [`ModelParallelizer`](#nemo_automodel-components-distributed-parallelizer-ModelParallelizer) | Single model-owned parallelization sidecar contract. |

### Functions

| Name                                                                                                                     | Description                                                           |
| ------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------- |
| [`_apply_model_parallelizer`](#nemo_automodel-components-distributed-model_parallelizer-_apply_model_parallelizer)       | Execute shared strategy dispatch for one model-owned parallelizer.    |
| [`_parallelize_ddp`](#nemo_automodel-components-distributed-model_parallelizer-_parallelize_ddp)                         | -                                                                     |
| [`_parallelize_fsdp2`](#nemo_automodel-components-distributed-model_parallelizer-_parallelize_fsdp2)                     | -                                                                     |
| [`_parallelize_megatron_fsdp`](#nemo_automodel-components-distributed-model_parallelizer-_parallelize_megatron_fsdp)     | -                                                                     |
| [`_parallelize_moe`](#nemo_automodel-components-distributed-model_parallelizer-_parallelize_moe)                         | -                                                                     |
| [`_parallelize_unsharded_fsdp2`](#nemo_automodel-components-distributed-model_parallelizer-_parallelize_unsharded_fsdp2) | -                                                                     |
| [`compile_parallelized_model`](#nemo_automodel-components-distributed-model_parallelizer-compile_parallelized_model)     | Compile FSDP2 layers after parallelization when requested.            |
| [`get_model_parallelizer`](#nemo_automodel-components-distributed-model_parallelizer-get_model_parallelizer)             | Return the class-owned sidecar, or the shared default implementation. |
| [`parallelize_model`](#nemo_automodel-components-distributed-model_parallelizer-parallelize_model)                       | Apply all requested parallelisms through the model-owned contract.    |

### Data

[`_DEFAULT_PARALLELIZER`](#nemo_automodel-components-distributed-model_parallelizer-_DEFAULT_PARALLELIZER)

### API

```python
class nemo_automodel.components.distributed.parallelizer.ModelParallelizer()
```

Single model-owned parallelization sidecar contract.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer._apply(
    model: torch.nn.Module,
    device_mesh: torch.distributed.device_mesh.DeviceMesh,
    mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None = None,
    offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
    sequence_parallel: bool = False,
    activation_checkpointing: bool = False,
    tp_shard_plan: typing.Union[typing.Dict[str, torch.distributed.tensor.parallel.ParallelStyle], str] | None = None,
    dp_replicate_mesh_name: str = 'dp_replicate',
    dp_shard_cp_mesh_name: str = 'dp_shard_cp',
    tp_mesh_name: str = 'tp',
    enable_async_tensor_parallel: bool = False,
    enable_compile: bool = False,
    enable_fsdp2_prefetch: bool = True,
    fsdp2_backward_prefetch_depth: int = 2,
    fsdp2_forward_prefetch_depth: int = 1,
    reshard_after_forward: bool | None = None,
    activation_checkpointing_scope: nemo_automodel.components.distributed.config.ActivationCheckpointingScope | None = 'all',
    frozen_multimodal_sharding: nemo_automodel.components.distributed.multimodal_fsdp.FrozenMultimodalSharding = 'root',
    reapply_trainability: collections.abc.Callable[[nn.Module], None] | None = None
) -> torch.nn.Module
```

Apply the shared dense FSDP2 implementation.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer._apply_fsdp_sharding(
    module: torch.nn.Module,
    mesh: torch.distributed.device_mesh.DeviceMesh,
    mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None,
    offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
    enable_fsdp2_prefetch: bool = True,
    fsdp2_backward_prefetch_depth: int = 2,
    fsdp2_forward_prefetch_depth: int = 1,
    reshard_after_forward: bool | None = None,
    frozen_multimodal_sharding: nemo_automodel.components.distributed.multimodal_fsdp.FrozenMultimodalSharding = 'root',
    ignored_multimodal_params: set[torch.nn.Parameter] | None = None
) -> None
```

Wrap the model's submodules into FSDP2 units.

Model-owned sidecar strategies deriving from this class override this
hook to change how parameters are grouped into FSDP units without
reimplementing the surrounding TP/AC/mixed-precision flow. Override
`_fully_shard_module` when a model needs a different primitive.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer._fully_shard_module(
    module: torch.nn.Module,
    kwargs = {}
) -> torch.nn.Module
```

Apply the FSDP2 primitive used by this model sidecar.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer._use_full_layer_activation_checkpointing(
    model: torch.nn.Module
) -> bool
```

Return whether this model safely opts into whole-layer checkpointing.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer._validate_tp_mesh(
    model: torch.nn.Module,
    tp_mesh: torch.distributed.device_mesh.DeviceMesh
) -> None
```

Validate the model's attention topology against its TP mesh.

```python
nemo_automodel.components.distributed.parallelizer.ModelParallelizer.parallelize(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> torch.nn.Module
```

Apply every requested parallelism and return the parallelized model.

```python
nemo_automodel.components.distributed.model_parallelizer._apply_model_parallelizer(
    parallelizer: nemo_automodel.components.distributed.parallelizer.ModelParallelizer,
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> torch.nn.Module
```

Execute shared strategy dispatch for one model-owned parallelizer.

```python
nemo_automodel.components.distributed.model_parallelizer._parallelize_ddp(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> torch.nn.Module
```

```python
nemo_automodel.components.distributed.model_parallelizer._parallelize_fsdp2(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext,
    parallelizer: nemo_automodel.components.distributed.parallelizer.ModelParallelizer
) -> torch.nn.Module
```

```python
nemo_automodel.components.distributed.model_parallelizer._parallelize_megatron_fsdp(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> torch.nn.Module
```

```python
nemo_automodel.components.distributed.model_parallelizer._parallelize_moe(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext,
    parallelizer: nemo_automodel.components.distributed.parallelizer.ModelParallelizer
) -> torch.nn.Module
```

```python
nemo_automodel.components.distributed.model_parallelizer._parallelize_unsharded_fsdp2(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext,
    parallelizer: nemo_automodel.components.distributed.parallelizer.ModelParallelizer
) -> torch.nn.Module
```

```python
nemo_automodel.components.distributed.model_parallelizer.compile_parallelized_model(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> None
```

Compile FSDP2 layers after parallelization when requested.

```python
nemo_automodel.components.distributed.model_parallelizer.get_model_parallelizer(
    model: torch.nn.Module
) -> nemo_automodel.components.distributed.parallelizer.ModelParallelizer
```

Return the class-owned sidecar, or the shared default implementation.

```python
nemo_automodel.components.distributed.model_parallelizer.parallelize_model(
    model: torch.nn.Module,
    mesh_context: nemo_automodel.components.distributed.mesh.MeshContext
) -> torch.nn.Module
```

Apply all requested parallelisms through the model-owned contract.

```python
nemo_automodel.components.distributed.model_parallelizer._DEFAULT_PARALLELIZER = ModelParallelizer()
```