nemo_automodel.recipes.llm.train_vispec
nemo_automodel.recipes.llm.train_vispec
ViSpec draft-training recipe for vision-language targets (stage 2).
ViSpec (arXiv:2509.15235) trains in two stages:
- Stage 1 — text only. EAGLE-1/2 training on a text corpus, with the
ranking term ViSpec adds to it (
rank_loss_weight); seeTrainVispecStage1Recipe. The resulting draft has no vision modules yet. - Stage 2 — vision aware. This recipe. It loads the stage-1 draft through
recipe_args.draft_init_from, adds the image adaptor and the global-image projection, and trains on image+text conversations whose assistant turns were regenerated by the target VLM (seenemo_automodel/components/speculative/regenerate_vlm.py).
The training loop, checkpointing, and resume behavior are inherited from
TrainEagle1Recipe; only the model/data construction and the per-batch
supervision differ.
Module Contents
Classes
Functions
Data
API
Bases: TrainEagle1Recipe
Recipe for ViSpec stage-2 (vision-aware) draft training.
Run the frozen VLM target and the ViSpec draft over one micro-batch.
Parameters:
Dataloader batch already on self.device, carrying
input_ids / attention_mask / loss_mask of shape
[1, sequence] plus the processor’s vision tensors.
Returns:
VispecStepMetrics for this micro-batch.
Initialize the shared draft weights from a stage-1 EAGLE-1/2 checkpoint.
The ViSpec-only modules (img_adaptor, img_fc) are absent from a
stage-1 checkpoint and keep their fresh initialization, which starts as
an identity pass-through, so stage 2 begins numerically equal to stage 1.
Parameters:
Path to a stage-1 checkpoint, either a directory
holding the consolidated export or a single .safetensors
file, or None to train the draft from scratch. A directory
is the practical form: the checkpointer names the export by
shard (model-00001-of-0000N.safetensors), so there is no
fixed file name to point at. Resolution is delegated to
load_hf_safetensors_state_dict, the same helper the EAGLE-3
recipe uses for its own warm start, so a sharded export with an
index file loads identically on both paths.
Return ViSpec’s two loss terms for logging.
Parameters:
The VispecStepMetrics returned by _compute_metrics.
Returns: dict[str, float]
Mapping of log-suffix to scalar value, logged as train/<key>.
Build the frozen VLM target, the ViSpec draft, data, optimizer, and trainer module.
Bases: TrainEagle1Recipe
Train ViSpec’s text-only EAGLE draft with a frozen VLM teacher.
ViSpec’s first stage does not consume images or construct ViSpec’s image adaptor. It does, however, use the intended VLM’s language tower as the teacher, so the draft learns the target’s text decoding distribution before stage 2 adds vision-aware training.
Build the frozen VLM teacher and text-only EAGLE draft for stage 1.
Build ViSpec’s train-only feature augmentation from the recipe config.
ViSpec enables the same sequence-scaled uniform noise in both stages, so
both call this. feature_noise_std: 0 disables it.
Raises:
ValueError: If the config still carries the old fixed-widthfeature_noisekey, which this no longer reads. Ignoring it would silently train at a different noise level than the config states.
Resolve a stage-1 checkpoint path to the directory holding its shards.
load_hf_safetensors_state_dict reads a directory of shards (honoring an
index file), but it returns None rather than raising when the directory
holds none, and it does not look one level down. The checkpointer writes the
export under <checkpoint>/model/consolidated, so accepting the checkpoint
root is what makes the config path usable, and a missing export has to fail
loudly: silently loading nothing would train stage 2 from a random draft.
Parameters:
A .safetensors file, or a directory holding the
consolidated export, or a directory one level above it.
Returns: str
The path to hand to load_hf_safetensors_state_dict.
Raises:
FileNotFoundError: If no.safetensorsfile is at or under the path.
Resolve the target’s image-placeholder token id.
Parameters:
The target VLM’s HuggingFace config.
The target’s AutoProcessor.
Returns: int
The token id that marks an image position in input_ids.
Seed ViSpec draft initialization and training-time stochasticity.
shuffle_seed predates the explicit seed knob and remains the
backwards-compatible default. Seeding must happen before the draft is
constructed because the frozen target is loaded from checkpoints while the
draft starts from random weights.
Entrypoint for the ViSpec stage selected by the recipe config.
Parameters:
Optional default YAML path. The command-line --config
argument overrides this value.
Raises:
ValueError: Ifrecipeis not a supported ViSpec recipe name.