nemo_automodel.components.speculative.dspark
nemo_automodel.components.speculative.dspark
DSpark speculative-decoding draft model and training objective.
A semi-autoregressive parallel drafter: a parallel backbone produces every position of a block in one pass, a lightweight serial Markov head injects intra-block token dependency, and a confidence head predicts per-position acceptance probability for scheduled verification.
Submodules
nemo_automodel.components.speculative.dspark._samplingnemo_automodel.components.speculative.dspark.commonnemo_automodel.components.speculative.dspark.confignemo_automodel.components.speculative.dspark.corenemo_automodel.components.speculative.dspark.draft_deepseek_v4nemo_automodel.components.speculative.dspark.draft_gemma4nemo_automodel.components.speculative.dspark.draft_glm_5_2nemo_automodel.components.speculative.dspark.draft_kimi_k3nemo_automodel.components.speculative.dspark.draft_minimax_m3nemo_automodel.components.speculative.dspark.draft_qwen3nemo_automodel.components.speculative.dspark.lossnemo_automodel.components.speculative.dspark.markov_headnemo_automodel.components.speculative.dspark.registrynemo_automodel.components.speculative.dspark.targetnemo_automodel.components.speculative.dspark.target_utils
Package Contents
Classes
Functions
API
Outputs for one DSpark training forward.
Shape symbols: batch_size: number of samples in the batch seq_len: source sequence length num_anchors: sampled anchor blocks per sample block_size: number of draft positions per anchor vocab_size: vocabulary size
The sampler keeps anchors whose first draft target is enabled by
loss_mask. Later slots are supervised only while they remain inside
seq_len and form a contiguous enabled prefix. Dummy anchors can still
appear when a sample has too few valid anchors; they are masked out by
block_keep_mask and eval_mask.
Bases: Qwen3PreTrainedModel
Keep the RoPE inv_freq buffer in fp32 across dtype casts.
model.to(bfloat16) (the training build path) would otherwise round
inv_freq to bf16 and dephase RoPE with absolute position, eroding
draft acceptance (see pin_rope_inv_freq_fp32).
Run one DSpark training forward.
Sequence packing (position_ids [B, S] per-document reset positions,
seq_lens [B, max_docs], doc_remaining [B, S]) keeps every block
inside its anchor’s document: the anchor’s first target must be in-document,
the block’s context prefix and supervision are restricted to that document,
and the draft’s RoPE uses the per-document positions.