> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.qwen3_5.parallelization

Model-owned distributed parallelization for dense Qwen3.5 models.

## Module Contents

### Classes

| Name                                                                                                             | Description                                                            |
| ---------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| [`Qwen3_5ModelParallelizer`](#nemo_automodel-components-models-qwen3_5-parallelization-Qwen3_5ModelParallelizer) | Keep mixed-dtype GatedDeltaNet parameters in dtype-uniform FSDP units. |

### Data

[`PARALLELIZER`](#nemo_automodel-components-models-qwen3_5-parallelization-PARALLELIZER)

[`__all__`](#nemo_automodel-components-models-qwen3_5-parallelization-__all__)

[`logger`](#nemo_automodel-components-models-qwen3_5-parallelization-logger)

### API

```python
class nemo_automodel.components.models.qwen3_5.parallelization.Qwen3_5ModelParallelizer()
```

**Bases:** `ModelParallelizer`

Keep mixed-dtype GatedDeltaNet parameters in dtype-uniform FSDP units.

**`_fp32_compute_module_names`** `tuple[str, ...] = ('_fp32_params',)`

---

```python
nemo_automodel.components.models.qwen3_5.parallelization.Qwen3_5ModelParallelizer._apply(
    model,
    device_mesh,
    dp_shard_cp_mesh_name = 'dp_shard_cp',
    kwargs = {}
)
```

Apply generic TP/AC/FSDP and install Qwen3.5's CP mesh.

```python
nemo_automodel.components.models.qwen3_5.parallelization.Qwen3_5ModelParallelizer._apply_fsdp_sharding(
    module: torch.nn.Module,
    mesh: torch.distributed.device_mesh.DeviceMesh,
    mp_policy: torch.distributed.fsdp.MixedPrecisionPolicy | None,
    offload_policy: torch.distributed.fsdp.OffloadPolicy | None = None,
    enable_fsdp2_prefetch: bool = True,
    fsdp2_backward_prefetch_depth: int = 2,
    fsdp2_forward_prefetch_depth: int = 1,
    reshard_after_forward: bool | None = None,
    frozen_multimodal_sharding: nemo_automodel.components.distributed.multimodal_fsdp.FrozenMultimodalSharding = 'root',
    ignored_multimodal_params: set[torch.nn.Parameter] | None = None
) -> None
```

Shard each decoder layer into dtype-uniform FSDP groups.

```python
nemo_automodel.components.models.qwen3_5.parallelization.PARALLELIZER = Qwen3_5ModelParallelizer()
```

```python
nemo_automodel.components.models.qwen3_5.parallelization.__all__ = ['PARALLELIZER']
```

```python
nemo_automodel.components.models.qwen3_5.parallelization.logger = logging.getLogger(__name__)
```