> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.kd_utils

Distributed topology and tensor transport helpers for KD recipes.

## Module Contents

### Classes

| Name                                                                          | Description                                                    |
| ----------------------------------------------------------------------------- | -------------------------------------------------------------- |
| [`KDDistributedSetups`](#nemo_automodel-recipes-kd_utils-KDDistributedSetups) | Student/teacher setups plus their global-rank assignments.     |
| [`KDMeshBridge`](#nemo_automodel-recipes-kd_utils-KDMeshBridge)               | Move batches and teacher logits between disjoint model meshes. |
| [`_ConfigLike`](#nemo_automodel-recipes-kd_utils-_ConfigLike)                 | Minimal recipe-config interface consumed by KD topology setup. |
| [`_Replica`](#nemo_automodel-recipes-kd_utils-_Replica)                       | -                                                              |
| [`_Route`](#nemo_automodel-recipes-kd_utils-_Route)                           | -                                                              |

### Functions

| Name                                                                                            | Description                                                                   |
| ----------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- |
| [`_mesh_size`](#nemo_automodel-recipes-kd_utils-_mesh_size)                                     | Return the explicitly requested mesh size for a separate KD model.            |
| [`_model_replicas`](#nemo_automodel-recipes-kd_utils-_model_replicas)                           | -                                                                             |
| [`_section_to_dict`](#nemo_automodel-recipes-kd_utils-_section_to_dict)                         | -                                                                             |
| [`_tree_from_spec`](#nemo_automodel-recipes-kd_utils-_tree_from_spec)                           | Rebuild a nested value from tensor leaves described by `_tree_spec`.          |
| [`_tree_spec`](#nemo_automodel-recipes-kd_utils-_tree_spec)                                     | Describe a nested value while extracting tensor leaves without copies.        |
| [`configure_kd_teacher_packing`](#nemo_automodel-recipes-kd_utils-configure_kd_teacher_packing) | Adapt NEAT-packed teachers and check both roles consume the same mask layout. |
| [`create_kd_distributed_setups`](#nemo_automodel-recipes-kd_utils-create_kd_distributed_setups) | Build shared or explicitly disjoint student and teacher setups.               |
| [`materialize_teacher_logits`](#nemo_automodel-recipes-kd_utils-materialize_teacher_logits)     | Reconstruct full teacher logits across TP and CP before mesh transport.       |

### Data

[`RUN_TEACHER`](#nemo_automodel-recipes-kd_utils-RUN_TEACHER)

[`STOP_TEACHER`](#nemo_automodel-recipes-kd_utils-STOP_TEACHER)

### API

```python
class nemo_automodel.recipes.kd_utils.KDDistributedSetups(
    student: nemo_automodel.components.distributed.config.DistributedSetup,
    teacher: nemo_automodel.components.distributed.config.DistributedSetup,
    student_ranks: tuple[int, ...],
    teacher_ranks: tuple[int, ...],
    separate: bool
)
```

Dataclass

Student/teacher setups plus their global-rank assignments.

**`separate`** `bool`

Whether the assignments are disjoint.

---

**`student`** `DistributedSetup`

Resolved student topology and policies.

---

**`student_ranks`** `tuple[int, ...]`

Ordered global ranks assigned to the student.

---

**`teacher`** `DistributedSetup`

Resolved teacher topology and policies.

---

**`teacher_ranks`** `tuple[int, ...]`

Ordered global ranks assigned to the teacher.

---

```python
class nemo_automodel.recipes.kd_utils.KDMeshBridge(
    setups: nemo_automodel.recipes.kd_utils.KDDistributedSetups,
    device: torch.device
)
```

Move batches and teacher logits between disjoint model meshes.

**Parameters:**

**`setups`** `KDDistributedSetups`

Resolved disjoint student and teacher setups.

---

**`device`** `torch.device`

Device used for transport tensors and collectives.

---

**`control_group`**

---

**`input_routes`** `list[list[_Route]] = []`

---

**`is_student`** `bool`

---

**`is_teacher`** `bool`

---

**`num_waves`**

---

**`output_routes`** `list[list[_Route]] = []`

---

**`rank`** `= dist.get_rank()`

---

**`student_group`** `= dist.new_group(ranks=(list(self.student_ranks)))`

---

**`student_ranks`** `= setups.student_ranks`

---

**`student_replicas`** `= _model_replicas(setups.student)`

---

**`teacher_group`** `= dist.new_group(ranks=(list(self.teacher_ranks)))`

---

**`teacher_ranks`** `= setups.teacher_ranks`

---

**`teacher_replicas`** `= _model_replicas(setups.teacher)`

---

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge._broadcast_tree(
    value: typing.Any,
    route: nemo_automodel.recipes.kd_utils._Route
) -> typing.Any
```

Broadcast a nested tensor tree along one route.

Tensor leaves may have arbitrary shapes and axis order.

**Parameters:**

**`value`** `Any`

Nested source value on `route.src` and `None` elsewhere.

---

**`route`** `_Route`

Source, membership, and process group for the broadcast.

---

**Returns:** `Any`

Reconstructed value on route members and `None` elsewhere.

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.broadcast_command(
    command: int | None = None
) -> int
```

Broadcast one worker command from the first student rank.

**Parameters:**

**`command`** `int | None` — default: None

Command supplied on student ranks. Teacher ranks pass
`None` while waiting for the broadcast.

---

**Returns:** `int`

Broadcast command value on every student and teacher rank.

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.match_student_vocab_shard(
    student_logits: torch.Tensor,
    teacher_logits: torch.Tensor
) -> torch.Tensor
```

staticmethod

Match full teacher logits to a TP-sharded student vocabulary.

**Parameters:**

**`student_logits`** `torch.Tensor`

Tensor of global shape
`[batch, sequence, vocab]` containing student logits. A
vocabulary-sharded `DTensor` has local shape
`[batch, sequence, local_vocab]` and `Shard(-1)` placement.

---

**`teacher_logits`** `torch.Tensor`

Replicated tensor of shape
`[batch, sequence, vocab]` containing teacher logits.

---

**Returns:** `torch.Tensor`

Replicated tensor of shape `[batch, sequence, vocab]` for a

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.move_to_device(
    value: typing.Any
) -> typing.Any
```

Move every tensor leaf in a nested value to the bridge device.

Tensor leaves may have arbitrary shapes and axis order.

**Parameters:**

**`value`** `Any`

Nested dictionaries, lists, tuples, scalar values, and tensor
leaves.

---

**Returns:** `Any`

Equivalent nested value whose tensors preserve shape and dtype on

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.send_batch(
    wave: int,
    batch: typing.Any | None
) -> typing.Any | None
```

Send one nested batch from each active student replica to a teacher.

Batch tensor leaves preserve their original shapes and axis order.

**Parameters:**

**`wave`** `int`

Routing wave index.

---

**`batch`** `Any | None`

Nested student batch on student ranks and `None` on teacher
ranks.

---

**Returns:** `Any | None`

Assigned nested batch on teacher ranks and `None` on student ranks.

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.send_logits(
    wave: int,
    logits: torch.Tensor | None
) -> torch.Tensor | None
```

Send full teacher logits back to the assigned student replica.

**Parameters:**

**`wave`** `int`

Routing wave index.

---

**`logits`** `torch.Tensor | None`

Tensor of shape `[batch, sequence, vocab]` containing full
replicated teacher logits on the teacher output rank, otherwise
`None`.

---

**Returns:** `torch.Tensor | None`

Replicated tensor of shape `[batch, sequence, vocab]` on ranks in

```python
nemo_automodel.recipes.kd_utils.KDMeshBridge.synchronize() -> None
```

Wait for both model roles without using the default process group.

```python
class nemo_automodel.recipes.kd_utils._ConfigLike()
```

Protocol

Minimal recipe-config interface consumed by KD topology setup.

```python
nemo_automodel.recipes.kd_utils._ConfigLike.get(
    key: str,
    default: typing.Any = None
) -> typing.Any
```

Return one configuration value or `default`.

```python
class nemo_automodel.recipes.kd_utils._Replica(
    ranks: tuple[int, ...],
    input_rank: int,
    output_rank: int
)
```

Dataclass

**`input_rank`** `int`

---

**`output_rank`** `int`

---

**`ranks`** `tuple[int, ...]`

---

```python
class nemo_automodel.recipes.kd_utils._Route(
    src: int,
    ranks: tuple[int, ...],
    group: torch.distributed.ProcessGroup
)
```

Dataclass

**`group`** `ProcessGroup`

---

**`ranks`** `tuple[int, ...]`

---

**`src`** `int`

---

```python
nemo_automodel.recipes.kd_utils._mesh_size(
    distributed_cfg: typing.Any,
    label: str
) -> int
```

Return the explicitly requested mesh size for a separate KD model.

```python
nemo_automodel.recipes.kd_utils._model_replicas(
    setup: nemo_automodel.components.distributed.config.DistributedSetup
) -> list[nemo_automodel.recipes.kd_utils._Replica]
```

```python
nemo_automodel.recipes.kd_utils._section_to_dict(
    section: typing.Any
) -> dict
```

```python
nemo_automodel.recipes.kd_utils._tree_from_spec(
    spec: typing.Any,
    tensors: list[torch.Tensor]
) -> typing.Any
```

Rebuild a nested value from tensor leaves described by `_tree_spec`.

Tensor leaves preserve the exact shapes and dtypes encoded in `spec` and
alias the corresponding entries in `tensors`.

**Parameters:**

**`spec`** `Any`

Metadata returned by `_tree_spec`.

---

**`tensors`** `list[torch.Tensor]`

Tensor leaves with the exact recorded shapes and dtypes.

---

**Returns:** `Any`

Reconstructed nested value whose tensor leaves alias `tensors`.

```python
nemo_automodel.recipes.kd_utils._tree_spec(
    value: typing.Any,
    tensors: list[torch.Tensor]
) -> typing.Any
```

Describe a nested value while extracting tensor leaves without copies.

Tensor leaves may have arbitrary rank and axis semantics; each leaf's exact
shape and dtype are recorded for allocation on receiving ranks.

**Parameters:**

**`value`** `Any`

Nested dictionaries, lists, tuples, scalar values, and tensor
leaves with arbitrary shape and axis order.

---

**`tensors`** `list[torch.Tensor]`

Output list populated with aliases of tensor leaves in traversal
order.

---

**Returns:** `Any`

Pickle-compatible nested metadata describing `value`.

```python
nemo_automodel.recipes.kd_utils.configure_kd_teacher_packing(
    teacher_parts: collections.abc.Sequence[torch.nn.Module],
    student_parts: collections.abc.Sequence[torch.nn.Module],
    control_group: torch.distributed.ProcessGroup | None = None
) -> None
```

Adapt NEAT-packed teachers and check both roles consume the same mask layout.

**Parameters:**

**`teacher_parts`** `Sequence[torch.nn.Module]`

Locally owned teacher stages, empty on student-only ranks.

---

**`student_parts`** `Sequence[torch.nn.Module]`

Locally owned student stages, empty on teacher-only ranks.

---

**`control_group`** `dist.ProcessGroup | None` — default: None

Shared group for separate-mesh KD; all its ranks must call.
Same-mesh KD passes None and performs only a local check.

---

**Raises:**

* `ValueError`: If teacher and student packed mask layouts disagree.

```python
nemo_automodel.recipes.kd_utils.create_kd_distributed_setups(
    cfg: nemo_automodel.recipes.kd_utils._ConfigLike,
    world_size: int
) -> nemo_automodel.recipes.kd_utils.KDDistributedSetups
```

Build shared or explicitly disjoint student and teacher setups.

**Parameters:**

**`cfg`** `_ConfigLike`

Recipe configuration containing `distributed` and optional
`teacher_distributed` sections.

---

**`world_size`** `int`

Total global process count available to both models.

---

**Returns:** `KDDistributedSetups`

Resolved model setups and their ordered global-rank assignments.

```python
nemo_automodel.recipes.kd_utils.materialize_teacher_logits(
    logits: torch.Tensor,
    device_mesh: 'DeviceMesh',
    sequence_length: int
) -> torch.Tensor
```

Reconstruct full teacher logits across TP and CP before mesh transport.

**Parameters:**

**`logits`** `torch.Tensor`

Tensor of global shape `[batch, sequence, vocab]` containing
teacher logits. It may be a vocabulary-sharded `DTensor` with
`Shard(-1)` placement and/or have a per-rank load-balanced
`local_sequence` extent under CP.

---

**`device_mesh`** `'DeviceMesh'`

Teacher mesh containing optional `tp` and `cp` axes.

---

**`sequence_length`** `int`

Unpadded global sequence length to retain.

---

**Returns:** `torch.Tensor`

Detached contiguous tensor of shape

```python
nemo_automodel.recipes.kd_utils.RUN_TEACHER = 1
```

```python
nemo_automodel.recipes.kd_utils.STOP_TEACHER = 0
```