nemo_rl.data.megatron_sft_packed#

Megatron-LM SFT packed JSONL preprocessing helpers.

Module Contents#

Classes#

_PromptConfig

MegatronSFTPackedDatumSpec

One offline-packed SFT row, pre-shifted for the direct Megatron-LM path.

Functions#

validate_megatron_sft_prompt_format

is_direct_packed_row

Return whether row came from the direct Megatron-LM prepacked path.

direct_packed_cp_granularity

Return the token multiple every direct-packed segment must satisfy.

split_megatron_sft_conversations

Split a packed row whenever a new system message begins.

resolve_megatron_sft_prompt_config

Resolve the tokenizer template, pad token, and assistant mask preset.

_normalize_token_ids

_tokenize_megatron_sft_conversation

count_conversation_tokens

Count tokens with the same prompt rendering and validation as the loader.

_resolve_pad_token_id

megatron_sft_packed_preprocessor

Build one direct tensor row with Megatron-LM SFTDataset semantics.

Data#

API#

nemo_rl.data.megatron_sft_packed.IGNORE_INDEX#

None

nemo_rl.data.megatron_sft_packed.IDENTITY_TEMPLATE#

“{% for message in messages %}{{ message[‘content’] }}{% endfor %}”

nemo_rl.data.megatron_sft_packed.NEMOTRON_H_ALIGNED_TEMPLATE = <Multiline-String>#
nemo_rl.data.megatron_sft_packed.NEMOTRON_NANO_V2_TEMPLATE = <Multiline-String>#
class nemo_rl.data.megatron_sft_packed._PromptConfig#
assistant_prefix_len: int#

None

pad_token: str | None#

None

chat_template: str#

None

has_bos: bool#

False

has_system_role: bool#

True

nemo_rl.data.megatron_sft_packed._PROMPT_CONFIGS#

None

nemo_rl.data.megatron_sft_packed.validate_megatron_sft_prompt_format(prompt_format: str) → None#
class nemo_rl.data.megatron_sft_packed.MegatronSFTPackedDatumSpec#

Bases: nemo_rl.data.interfaces.DatumSpec

One offline-packed SFT row, pre-shifted for the direct Megatron-LM path.

A .jsonl.packed record is re-tokenized at load time into a single row of exactly max_seq_length target-aligned positions: input_ids is the pack without its last token and target_ids is the pack without its first, so nothing downstream shifts again.

.. attribute:: input_ids

[max_seq_length] token ids, pack[:-1].

.. attribute:: target_ids

[max_seq_length] labels, pack[1:]. Prompt positions carry IGNORE_INDEX; trailing padding carries the pad id.

.. attribute:: token_mask

[max_seq_length] float mask, 0.0 wherever target_ids is padding or IGNORE_INDEX. This is not cosmetic: mcore’s fused cross-entropy clamps IGNORE_INDEX before the vocab gather and forces predicted_logits = 0.0, so an ignored position still emits log(sum_exp_logits) - a large finite loss with a non-zero gradient. token_mask is the only thing that zeroes both the loss and the gradient at those positions.

.. attribute:: position_ids

[max_seq_length] positions that restart at 0 on every packed segment boundary.

.. attribute:: packed_cu_seqlens

[num_segments + 1] int32 cumulative segment boundaries, from 0 through the packed row length. Its presence is the single marker that a row took this path; see

Func:

is_direct_packed_row.

.. attribute:: packed_max_seqlen

Longest packed segment, used to build PackedSeqParams.

.. attribute:: packed_context_parallel_size

Context-parallel size the row was padded for. Every segment length is a multiple of

Func:

direct_packed_cp_granularity because mcore shards packed rows with per-document zigzag.

Initialization

Initialize self. See help(type(self)) for accurate signature.

input_ids: torch.Tensor#

None

target_ids: torch.Tensor#

None

token_mask: torch.Tensor#

None

position_ids: torch.Tensor#

None

packed_cu_seqlens: torch.Tensor#

None

packed_max_seqlen: int#

None

packed_context_parallel_size: int#

None

nemo_rl.data.megatron_sft_packed.MEGATRON_SFT_PACKED_FIELDS#

None

nemo_rl.data.megatron_sft_packed.MEGATRON_SFT_PACKED_BATCH_FIELDS#

None

nemo_rl.data.megatron_sft_packed.is_direct_packed_row(
row: collections.abc.Mapping[str, Any],
) → bool#

Return whether row came from the direct Megatron-LM prepacked path.

packed_cu_seqlens is the one authoritative marker. Every other field in

Data:

MEGATRON_SFT_PACKED_FIELDS is emitted alongside it by

Func:

megatron_sft_packed_preprocessor, so testing any other key answers a different question and the answers drift apart.

nemo_rl.data.megatron_sft_packed.direct_packed_cp_granularity(context_parallel_size: int) → int#

Return the token multiple every direct-packed segment must satisfy.

mcore shards packed rows with per-document zigzag, which cuts each document into 2 * cp_size chunks, so a segment length that is not a multiple of this value cannot be sharded.

nemo_rl.data.megatron_sft_packed.split_megatron_sft_conversations(
merged_messages: list[dict[str, Any]],
) → list[list[dict[str, Any]]]#

Split a packed row whenever a new system message begins.

nemo_rl.data.megatron_sft_packed.resolve_megatron_sft_prompt_config(
prompt_format: str,
override_pad_token: str | None = None,
) → nemo_rl.data.megatron_sft_packed._PromptConfig#

Resolve the tokenizer template, pad token, and assistant mask preset.

nemo_rl.data.megatron_sft_packed._normalize_token_ids(token_ids: Any) → list[int]#
nemo_rl.data.megatron_sft_packed._tokenize_megatron_sft_conversation(
conversation: list[dict[str, Any]],
tokenizer: Any,
prompt_format: str,
prompt_config: nemo_rl.data.megatron_sft_packed._PromptConfig,
) → tuple[list[int], list[int]]#
nemo_rl.data.megatron_sft_packed.count_conversation_tokens(
conversation: list[dict[str, Any]],
tokenizer: Any,
prompt_format: str,
) → int#

Count tokens with the same prompt rendering and validation as the loader.

nemo_rl.data.megatron_sft_packed._resolve_pad_token_id(
tokenizer: Any,
prompt_config: nemo_rl.data.megatron_sft_packed._PromptConfig,
) → int#
nemo_rl.data.megatron_sft_packed.megatron_sft_packed_preprocessor(
datum_dict: dict[str, Any],
task_data_spec: nemo_rl.data.interfaces.TaskDataSpec,
tokenizer: Any,
max_seq_length: int | None,
idx: int,
*,
prompt_format: str,
context_parallel_size: int,
pad_token: str | None = None,
**_unused_kwargs: Any,
) → nemo_rl.data.megatron_sft_packed.MegatronSFTPackedDatumSpec#

Build one direct tensor row with Megatron-LM SFTDataset semantics.