> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.nemotron_v3.parallelization

Model-owned distributed parallelization for Nemotron-H and Nemotron-V3.

## Module Contents

### Classes

| Name                                                                                                                     | Description                                                 |
| ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------- |
| [`NemotronHModelParallelizer`](#nemo_automodel-components-models-nemotron_v3-parallelization-NemotronHModelParallelizer) | Apply Nemotron-H's specialized TP, CP, AC, and FSDP policy. |

### Functions

| Name                                                                                               | Description                                                  |
| -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
| [`_decoder_blocks`](#nemo_automodel-components-models-nemotron_v3-parallelization-_decoder_blocks) | Return the mutable decoder container and its ordered blocks. |

### Data

[`PARALLELIZER`](#nemo_automodel-components-models-nemotron_v3-parallelization-PARALLELIZER)

[`__all__`](#nemo_automodel-components-models-nemotron_v3-parallelization-__all__)

[`logger`](#nemo_automodel-components-models-nemotron_v3-parallelization-logger)

### API

```python
class nemo_automodel.components.models.nemotron_v3.parallelization.NemotronHModelParallelizer()
```

**Bases:** `ModelParallelizer`

Apply Nemotron-H's specialized TP, CP, AC, and FSDP policy.

```python
nemo_automodel.components.models.nemotron_v3.parallelization.NemotronHModelParallelizer._apply(
    model: torch.nn.Module,
    device_mesh: torch.distributed.device_mesh.DeviceMesh,
    mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None = None,
    offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
    sequence_parallel: bool = False,
    activation_checkpointing: bool = False,
    tp_shard_plan: typing.Union[typing.Dict[str, torch.distributed.tensor.parallel.ParallelStyle], str] | None = None,
    dp_replicate_mesh_name: str = 'dp_replicate',
    dp_shard_cp_mesh_name: str = 'dp_shard_cp',
    tp_mesh_name: str = 'tp',
    reshard_after_forward: bool | None = None,
    reapply_trainability: collections.abc.Callable[[nn.Module], None] | None = None,
    kwargs = {}
) -> torch.nn.Module
```

Apply every requested parallelism to a Nemotron-H model.

```python
nemo_automodel.components.models.nemotron_v3.parallelization._decoder_blocks(
    model: torch.nn.Module
) -> tuple[torch.nn.Module, list[torch.nn.Module]]
```

Return the mutable decoder container and its ordered blocks.

```python
nemo_automodel.components.models.nemotron_v3.parallelization.PARALLELIZER = NemotronHModelParallelizer()
```

```python
nemo_automodel.components.models.nemotron_v3.parallelization.__all__ = ['PARALLELIZER']
```

```python
nemo_automodel.components.models.nemotron_v3.parallelization.logger = logging.getLogger(__name__)
```