nemo_rl.data.megatron_sft_packed#
Megatron-LM SFT packed JSONL preprocessing helpers.
Module Contents#
Classes#
One offline-packed SFT row, pre-shifted for the direct Megatron-LM path. |
Functions#
Return whether |
|
Return the token multiple every direct-packed segment must satisfy. |
|
Split a packed row whenever a new system message begins. |
|
Resolve the tokenizer template, pad token, and assistant mask preset. |
|
Count tokens with the same prompt rendering and validation as the loader. |
|
Build one direct tensor row with Megatron-LM |
Data#
API#
- nemo_rl.data.megatron_sft_packed.IGNORE_INDEX#
None
- nemo_rl.data.megatron_sft_packed.IDENTITY_TEMPLATE#
“{% for message in messages %}{{ message[‘content’] }}{% endfor %}”
- nemo_rl.data.megatron_sft_packed.NEMOTRON_H_ALIGNED_TEMPLATE = <Multiline-String>#
- nemo_rl.data.megatron_sft_packed.NEMOTRON_NANO_V2_TEMPLATE = <Multiline-String>#
- class nemo_rl.data.megatron_sft_packed._PromptConfig#
- assistant_prefix_len: int#
None
- pad_token: str | None#
None
- chat_template: str#
None
- has_bos: bool#
False
- has_system_role: bool#
True
- nemo_rl.data.megatron_sft_packed._PROMPT_CONFIGS#
None
- nemo_rl.data.megatron_sft_packed.validate_megatron_sft_prompt_format(prompt_format: str) None#
- class nemo_rl.data.megatron_sft_packed.MegatronSFTPackedDatumSpec#
Bases:
nemo_rl.data.interfaces.DatumSpecOne offline-packed SFT row, pre-shifted for the direct Megatron-LM path.
A
.jsonl.packedrecord is re-tokenized at load time into a single row of exactlymax_seq_lengthtarget-aligned positions:input_idsis the pack without its last token andtarget_idsis the pack without its first, so nothing downstream shifts again... attribute:: input_ids
[max_seq_length]token ids,pack[:-1]... attribute:: target_ids
[max_seq_length]labels,pack[1:]. Prompt positions carryIGNORE_INDEX; trailing padding carries the pad id... attribute:: token_mask
[max_seq_length]float mask,0.0wherevertarget_idsis padding orIGNORE_INDEX. This is not cosmetic: mcore’s fused cross-entropy clampsIGNORE_INDEXbefore the vocab gather and forcespredicted_logits = 0.0, so an ignored position still emitslog(sum_exp_logits)- a large finite loss with a non-zero gradient.token_maskis the only thing that zeroes both the loss and the gradient at those positions... attribute:: position_ids
[max_seq_length]positions that restart at 0 on every packed segment boundary... attribute:: packed_cu_seqlens
[num_segments + 1]int32 cumulative segment boundaries, from 0 through the packed row length. Its presence is the single marker that a row took this path; see- Func:
is_direct_packed_row.
.. attribute:: packed_max_seqlen
Longest packed segment, used to build
PackedSeqParams... attribute:: packed_context_parallel_size
Context-parallel size the row was padded for. Every segment length is a multiple of
- Func:
direct_packed_cp_granularitybecause mcore shards packed rows with per-document zigzag.
Initialization
Initialize self. See help(type(self)) for accurate signature.
- input_ids: torch.Tensor#
None
- target_ids: torch.Tensor#
None
- token_mask: torch.Tensor#
None
- position_ids: torch.Tensor#
None
- packed_cu_seqlens: torch.Tensor#
None
- packed_max_seqlen: int#
None
- packed_context_parallel_size: int#
None
- nemo_rl.data.megatron_sft_packed.MEGATRON_SFT_PACKED_FIELDS#
None
- nemo_rl.data.megatron_sft_packed.MEGATRON_SFT_PACKED_BATCH_FIELDS#
None
- nemo_rl.data.megatron_sft_packed.is_direct_packed_row(
- row: collections.abc.Mapping[str, Any],
Return whether
rowcame from the direct Megatron-LM prepacked path.packed_cu_seqlensis the one authoritative marker. Every other field in- Data:
MEGATRON_SFT_PACKED_FIELDSis emitted alongside it by- Func:
megatron_sft_packed_preprocessor, so testing any other key answers a different question and the answers drift apart.
- nemo_rl.data.megatron_sft_packed.direct_packed_cp_granularity(context_parallel_size: int) int#
Return the token multiple every direct-packed segment must satisfy.
mcore shards packed rows with per-document zigzag, which cuts each document into
2 * cp_sizechunks, so a segment length that is not a multiple of this value cannot be sharded.
- nemo_rl.data.megatron_sft_packed.split_megatron_sft_conversations(
- merged_messages: list[dict[str, Any]],
Split a packed row whenever a new system message begins.
- nemo_rl.data.megatron_sft_packed.resolve_megatron_sft_prompt_config(
- prompt_format: str,
- override_pad_token: str | None = None,
Resolve the tokenizer template, pad token, and assistant mask preset.
- nemo_rl.data.megatron_sft_packed._normalize_token_ids(token_ids: Any) list[int]#
- nemo_rl.data.megatron_sft_packed._tokenize_megatron_sft_conversation(
- conversation: list[dict[str, Any]],
- tokenizer: Any,
- prompt_format: str,
- prompt_config: nemo_rl.data.megatron_sft_packed._PromptConfig,
- nemo_rl.data.megatron_sft_packed.count_conversation_tokens(
- conversation: list[dict[str, Any]],
- tokenizer: Any,
- prompt_format: str,
Count tokens with the same prompt rendering and validation as the loader.
- nemo_rl.data.megatron_sft_packed._resolve_pad_token_id(
- tokenizer: Any,
- prompt_config: nemo_rl.data.megatron_sft_packed._PromptConfig,
- nemo_rl.data.megatron_sft_packed.megatron_sft_packed_preprocessor(
- datum_dict: dict[str, Any],
- task_data_spec: nemo_rl.data.interfaces.TaskDataSpec,
- tokenizer: Any,
- max_seq_length: int | None,
- idx: int,
- *,
- prompt_format: str,
- context_parallel_size: int,
- pad_token: str | None = None,
- **_unused_kwargs: Any,
Build one direct tensor row with Megatron-LM
SFTDatasetsemantics.