> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.datasets.packing

Dataset-owned packed-sequence construction contracts and helpers.

## Module Contents

### Classes

| Name                                                                                                           | Description                                                           |
| -------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| [`PackedSequenceContract`](#nemo_automodel-components-datasets-packing-PackedSequenceContract)                 | Structural model contract consumed while collating packed data.       |
| [`PackedSequenceMetadata`](#nemo_automodel-components-datasets-packing-PackedSequenceMetadata)                 | Batch-major metadata that remains valid after microbatch splitting.   |
| [`_DefaultPackedSequenceContract`](#nemo_automodel-components-datasets-packing-_DefaultPackedSequenceContract) | Block-causal packing defaults for callers without model requirements. |

### Functions

| Name                                                                                                           | Description                                                       |
| -------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- |
| [`build_packed_sequence_metadata`](#nemo_automodel-components-datasets-packing-build_packed_sequence_metadata) | Build batch-major metadata for a padded indexed packing mask.     |
| [`get_seqlens_in_batch`](#nemo_automodel-components-datasets-packing-get_seqlens_in_batch)                     | Extract document lengths from an indexed packed-sequence mask.    |
| [`get_unpad_data`](#nemo_automodel-components-datasets-packing-get_unpad_data)                                 | Build varlen metadata for an indexed or binary attention mask.    |
| [`resolve_packing_contract`](#nemo_automodel-components-datasets-packing-resolve_packing_contract)             | Translate the deprecated attention keyword to a packing contract. |

### Data

[`DEFAULT_PACKED_SEQUENCE_CONTRACT`](#nemo_automodel-components-datasets-packing-DEFAULT_PACKED_SEQUENCE_CONTRACT)

[`PackedMaskType`](#nemo_automodel-components-datasets-packing-PackedMaskType)

[`_LEGACY_FLASH_ATTENTION_IMPLEMENTATIONS`](#nemo_automodel-components-datasets-packing-_LEGACY_FLASH_ATTENTION_IMPLEMENTATIONS)

### API

```python
class nemo_automodel.components.datasets.packing.PackedSequenceContract()
```

Protocol

Structural model contract consumed while collating packed data.

**`packed_mask_type`** `PackedMaskType`

Packed attention-mask representation required by the model.

---

**`requires_packed_sequence_metadata`** `bool`

Whether the model consumes flat token indices and cumulative lengths.

---

```python
class nemo_automodel.components.datasets.packing.PackedSequenceMetadata
```

**Bases:** `typing.TypedDict`

Batch-major metadata that remains valid after microbatch splitting.

**`cu_seqlens`** `Tensor`

---

**`max_seqlen`** `int`

---

**`packed_token_indices`** `Tensor`

---

```python
class nemo_automodel.components.datasets.packing._DefaultPackedSequenceContract(
    packed_mask_type: nemo_automodel.components.datasets.packing.PackedMaskType = 'block_causal',
    requires_packed_sequence_metadata: bool = False
)
```

Dataclass

Block-causal packing defaults for callers without model requirements.

**`packed_mask_type`** `PackedMaskType = 'block_causal'`

---

**`requires_packed_sequence_metadata`** `bool = False`

---

```python
nemo_automodel.components.datasets.packing.build_packed_sequence_metadata(
    attention_mask: torch.Tensor
) -> nemo_automodel.components.datasets.packing.PackedSequenceMetadata
```

Build batch-major metadata for a padded indexed packing mask.

**Parameters:**

**`attention_mask`** `torch.Tensor`

Integer tensor of shape \[batch, sequence] containing
1-based document IDs and zero-valued padding.

---

**Returns:** `PackedSequenceMetadata`

Metadata containing row-local `packed_token_indices` of shape

```python
nemo_automodel.components.datasets.packing.get_seqlens_in_batch(
    attention_mask: torch.Tensor
) -> torch.Tensor
```

Extract document lengths from an indexed packed-sequence mask.

**Parameters:**

**`attention_mask`** `torch.Tensor`

Integer tensor of shape \[batch, sequence]. Each nonzero
value is a 1-based document index local to its batch row; zero marks
padding.

---

**Returns:** `torch.Tensor`

Tensor of shape \[documents] containing nonzero document lengths in

```python
nemo_automodel.components.datasets.packing.get_unpad_data(
    attention_mask: torch.Tensor
) -> tuple[torch.Tensor, torch.Tensor, int]
```

Build varlen metadata for an indexed or binary attention mask.

Indexed masks treat every distinct positive document index in each batch
row as a separate sequence. Binary masks treat every nonempty batch row as
one sequence. Padding tokens are omitted from the flattened token stream.

**Parameters:**

**`attention_mask`** `torch.Tensor`

Integer or boolean tensor of shape \[batch, sequence].
Positive values identify valid tokens and zero marks padding.

---

**Returns:** `torch.Tensor`

A tuple containing `indices` of shape \[tokens] into the flattened

**Raises:**

* `ValueError`: If the mask is not rank two or contains no valid tokens.

```python
nemo_automodel.components.datasets.packing.resolve_packing_contract(
    packing: nemo_automodel.components.datasets.packing.PackedSequenceContract,
    attn_implementation: str | None
) -> nemo_automodel.components.datasets.packing.PackedSequenceContract
```

Translate the deprecated attention keyword to a packing contract.

**Parameters:**

**`packing`** `PackedSequenceContract`

Explicit structural packing contract. It takes precedence when
both migration surfaces are supplied.

---

**`attn_implementation`** `str | None`

Deprecated attention-backend name, or `None`.

---

**Returns:** `PackedSequenceContract`

The explicit contract, or a compatibility contract matching the legacy

```python
nemo_automodel.components.datasets.packing.DEFAULT_PACKED_SEQUENCE_CONTRACT: Final[PackedSequenceContract] = _DefaultPackedSequenceContract()
```

```python
nemo_automodel.components.datasets.packing.PackedMaskType = Literal['block_causal', 'document_ids', 'flash_varlen']
```

```python
nemo_automodel.components.datasets.packing._LEGACY_FLASH_ATTENTION_IMPLEMENTATIONS = frozenset({'flash_attention_2', 'flash_attention_3', 'flash_attention_4'})
```