Packed Sequences#
Packed sequences are a fine-tuning technique that reduces padding waste by concatenating multiple examples into one pack while preserving sequence boundaries for attention. In Megatron Bridge, this is primarily a supervised fine-tuning and PEFT optimization rather than a general pretraining feature.
This page is the stable overview for what packed sequences are, when to use them, and which constraints are durable. For operational setup, code anchors, and verification commands, see skills/nemo-mbridge-perf-sequence-packing/SKILL.md.
What It Is#
Fine-tuning datasets often contain examples with highly variable lengths. When those examples are batched conventionally, many tokens in each batch are just padding. Packed sequences reduce that waste by building longer packs from multiple examples and carrying boundary metadata into the attention path.
In Bridge today, there are two distinct packing paths plus long-context enablement through context parallelism:
Path |
Use case |
Key config |
|---|---|---|
Offline packed SFT |
Text-only finetuning |
|
Direct-HF/VLM in-batch packing |
Direct Hugging Face and supported VLM finetuning |
|
Long-context (CP) |
Pretrain / finetune at 16K-128K+ |
|
These are related but they are not the same knob. Offline packed SFT and Direct-HF/VLM in-batch packing solve padding waste; long-context training primarily addresses activation memory and communication tradeoffs at larger sequence lengths.
The shared implementation lives under megatron.bridge.data.packing: offline
GPT SFT materialization, packed Parquet runtime datasets, bin-packing
algorithms, and collate-time THD packing each have separate modules. Ordinary
non-packed padding remains in megatron.bridge.data.collators. Use
scripts/training/prepare_gpt_sft_packed_data.py when packed GPT SFT artifacts
should be prepared before launching training.
When to Use It#
Packed sequences are a good fit when all of the following are true:
you are doing SFT, PEFT, or supported VLM finetuning using one of the two packing paths above
your examples have variable lengths and padding waste is significant
you can tolerate the micro-batch constraints of packed training
Packed sequences are usually not the right answer when:
you are doing standard Megatron-style pretraining, which already concatenates documents during sampling
you want long-context training in general, where context parallelism is often the main technique
your model family or recipe explicitly opts out of packed-sequence support
Choosing the Offline Pack Length#
For text-only LLM SFT and PEFT, use 8192 as the first offline-pack target when the model context limit, memory, and recipe support allow it. Compare candidate lengths at the same token slots per optimizer step:
token_slots_per_step = packed_sequence_size * global_batch_size
Thus 2K/GBS32, 4K/GBS16, and 8K/GBS8 each retain 65,536 token slots per step. A longer pack can contain more source examples in the physical MBS1 row and reduce gradient accumulation and launch overhead. It also consumes more activation memory and can encounter fixed-width kernel constraints, so measure the candidates rather than treating 8K as unconditional.
Offline packing requires MBS1. The selected GBS must be divisible by and no smaller than data parallel size; an 8K/GBS8 run therefore requires DP no larger than 8. Keep model, dataset, and packed sequence lengths equal, write changed packing configurations to a fresh output root, and verify the resolved post-setup configuration. Changing pack length can alter truncation and pack membership even when token slots stay constant, so rerun loss and stability checks before replacing existing verification evidence.
Derive the internal sequence alignment from the resolved topology for both SFT and PEFT:
cp_multiple = 2 * CP if CP > 1 else 1
sp_multiple = CP * TP if sequence parallelism is enabled and TP > 1 else 1
pad_seq_to_mult = lcm(cp_multiple, sp_multiple)
For example, TP1/CP1 with SP disabled uses 1, while TP4/CP1 with SP enabled uses 4. The difference comes from execution topology, not from whether the trainable set is full SFT or PEFT. Pin the derived value explicitly and rebuild the packed output after changing topology because the alignment changes pack membership.
Keep this internal alignment separate from fixed final pack width.
pad_to_max_length=true is needed when a dispatcher or kernel requires a
static width, such as a HybridEP combine kernel with a fixed token chunk, or
when using CUDA graphs. CUDA graphs additionally require
pad_cu_seqlens=true and packing metadata. Ordinary eager offline packing
does not universally require fixed-width padding.
Stable Constraints#
The durable constraints for packed sequences in Bridge are:
offline packed SFT requires configured
micro_batch_size == 1Direct-HF/VLM in-batch packing requires configured
micro_batch_size > 1; collation flattens those input rows into one physical THD batch rowwhen context parallelism is used, sequence length must satisfy the standard CP divisibility constraints
Direct-HF sequence length must also satisfy the LCM of the training and evaluation CP constraints and
CP * TPwhen sequence parallelism is enabledfor fine-tuning with CP enabled, per-token loss behavior and reduction settings matter
Megatron Bridge automatically enables safe uneven-input padding for eager HybridEP configs; this pads only to the group-wide aligned maximum before dispatch and trims the padding after combine
CUDA-graph-friendly packed metadata requires additional padding constraints
Model-family support is not universal. Some families and recipe paths explicitly opt out of packed sequences or related packing modes.
HybridEP CUDA-graph configs preserve their explicit uneven-input setting because the safety path performs a host scalar synchronization that is not capture-safe. They must provide equal per-rank dispatch shapes. Disable CUDA graphs when packed runtime token counts can differ so Bridge can enable safe padding.
Relationship to Long-Sequence Training#
Packed sequences and long-sequence training are often mentioned together because both affect sequence layout and memory behavior, but they solve different problems:
packed sequences mainly reduce padding waste in fine-tuning datasets
long-sequence training mainly addresses activation memory and communication tradeoffs at larger sequence lengths
For long-sequence training guidance, see:
docs/performance-guide.mddocs/training/hierarchical-context-parallel.md
Practical Caveats#
The most stable caveats to remember are:
Packed-sequence support is recipe- and model-family-specific.
Fine-tuning sequence packing should not be assumed to work with every other training feature.
Setting a distinct evaluation CP only reserves compatible data shapes; activating it requires decentralized process groups and caller-managed eval groups. The eval-CP example demonstrates topology plumbing, not a complete real-data recipe; validation sharding and batch math must use the eval DP.
Packed sequences improve efficiency primarily by reducing padding waste, not by replacing long-context parallelism or memory-planning techniques.