> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.llm.train_vispec

ViSpec draft-training recipe for vision-language targets (stage 2).

ViSpec (arXiv:2509.15235) trains in two stages:

1. **Stage 1 -- text only.** EAGLE-1/2 training on a text corpus, with the
   ranking term ViSpec adds to it (`rank_loss_weight`); see
   `TrainVispecStage1Recipe`. The resulting draft has no vision modules yet.
2. **Stage 2 -- vision aware.** This recipe. It loads the stage-1 draft through
   `recipe_args.draft_init_from`, adds the image adaptor and the global-image
   projection, and trains on image+text conversations whose assistant turns were
   regenerated by the target VLM (see
   `nemo_automodel/components/speculative/regenerate_vlm.py`).

The training loop, checkpointing, and resume behavior are inherited from
`TrainEagle1Recipe`; only the model/data construction and the per-batch
supervision differ.

## Module Contents

### Classes

| Name                                                                                          | Description                                                     |
| --------------------------------------------------------------------------------------------- | --------------------------------------------------------------- |
| [`TrainVispecRecipe`](#nemo_automodel-recipes-llm-train_vispec-TrainVispecRecipe)             | Recipe for ViSpec stage-2 (vision-aware) draft training.        |
| [`TrainVispecStage1Recipe`](#nemo_automodel-recipes-llm-train_vispec-TrainVispecStage1Recipe) | Train ViSpec's text-only EAGLE draft with a frozen VLM teacher. |

### Functions

| Name                                                                                                  | Description                                                            |
| ----------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| [`_build_feature_noise_config`](#nemo_automodel-recipes-llm-train_vispec-_build_feature_noise_config) | Build ViSpec's train-only feature augmentation from the recipe config. |
| [`_resolve_draft_export`](#nemo_automodel-recipes-llm-train_vispec-_resolve_draft_export)             | Resolve a stage-1 checkpoint path to the directory holding its shards. |
| [`_resolve_image_token_id`](#nemo_automodel-recipes-llm-train_vispec-_resolve_image_token_id)         | Resolve the target's image-placeholder token id.                       |
| [`_seed_draft_initialization`](#nemo_automodel-recipes-llm-train_vispec-_seed_draft_initialization)   | Seed ViSpec draft initialization and training-time stochasticity.      |
| [`main`](#nemo_automodel-recipes-llm-train_vispec-main)                                               | Entrypoint for the ViSpec stage selected by the recipe config.         |

### Data

[`logger`](#nemo_automodel-recipes-llm-train_vispec-logger)

### API

```python
class nemo_automodel.recipes.llm.train_vispec.TrainVispecRecipe()
```

**Bases:** [TrainEagle1Recipe](/nemo-automodel/nemo_automodel/recipes/llm/train_eagle1#nemo_automodel-recipes-llm-train_eagle1-TrainEagle1Recipe)

Recipe for ViSpec stage-2 (vision-aware) draft training.

```python
nemo_automodel.recipes.llm.train_vispec.TrainVispecRecipe._compute_metrics(
    batch: dict[str, torch.Tensor]
)
```

Run the frozen VLM target and the ViSpec draft over one micro-batch.

**Parameters:**

**`batch`** `dict[str, torch.Tensor]`

Dataloader batch already on `self.device`, carrying
`input_ids` / `attention_mask` / `loss_mask` of shape
\[1, sequence] plus the processor's vision tensors.

---

**Returns:**

VispecStepMetrics for this micro-batch.

```python
nemo_automodel.recipes.llm.train_vispec.TrainVispecRecipe._load_stage1_draft(
    draft_init_from: str | None
) -> None
```

Initialize the shared draft weights from a stage-1 EAGLE-1/2 checkpoint.

The ViSpec-only modules (`img_adaptor`, `img_fc`) are absent from a
stage-1 checkpoint and keep their fresh initialization, which starts as
an identity pass-through, so stage 2 begins numerically equal to stage 1.

**Parameters:**

**`draft_init_from`** `str | None`

Path to a stage-1 checkpoint, either a directory
holding the consolidated export or a single `.safetensors`
file, or `None` to train the draft from scratch. A directory
is the practical form: the checkpointer names the export by
shard (`model-00001-of-0000N.safetensors`), so there is no
fixed file name to point at. Resolution is delegated to
`load_hf_safetensors_state_dict`, the same helper the EAGLE-3
recipe uses for its own warm start, so a sharded export with an
index file loads identically on both paths.

---

```python
nemo_automodel.recipes.llm.train_vispec.TrainVispecRecipe._loss_components(
    metrics
) -> dict[str, float]
```

Return ViSpec's two loss terms for logging.

**Parameters:**

**`metrics`**

The `VispecStepMetrics` returned by `_compute_metrics`.

---

**Returns:** `dict[str, float]`

Mapping of log-suffix to scalar value, logged as `train/&lt;key&gt;`.

```python
nemo_automodel.recipes.llm.train_vispec.TrainVispecRecipe.setup()
```

Build the frozen VLM target, the ViSpec draft, data, optimizer, and trainer module.

```python
class nemo_automodel.recipes.llm.train_vispec.TrainVispecStage1Recipe()
```

**Bases:** [TrainEagle1Recipe](/nemo-automodel/nemo_automodel/recipes/llm/train_eagle1#nemo_automodel-recipes-llm-train_eagle1-TrainEagle1Recipe)

Train ViSpec's text-only EAGLE draft with a frozen VLM teacher.

ViSpec's first stage does not consume images or construct ViSpec's image
adaptor. It does, however, use the intended VLM's language tower as the
teacher, so the draft learns the target's text decoding distribution before
stage 2 adds vision-aware training.

```python
nemo_automodel.recipes.llm.train_vispec.TrainVispecStage1Recipe.setup()
```

Build the frozen VLM teacher and text-only EAGLE draft for stage 1.

```python
nemo_automodel.recipes.llm.train_vispec._build_feature_noise_config(
    recipe_cfg
) -> nemo_automodel.components.speculative.eagle.core_v12.FeatureNoiseConfig | None
```

Build ViSpec's train-only feature augmentation from the recipe config.

ViSpec enables the same sequence-scaled uniform noise in both stages, so
both call this. `feature_noise_std: 0` disables it.

**Raises:**

* `ValueError`: If the config still carries the old fixed-width
  `feature_noise` key, which this no longer reads. Ignoring it would
  silently train at a different noise level than the config states.

```python
nemo_automodel.recipes.llm.train_vispec._resolve_draft_export(
    draft_init_from: str
) -> str
```

Resolve a stage-1 checkpoint path to the directory holding its shards.

`load_hf_safetensors_state_dict` reads a directory of shards (honoring an
index file), but it returns `None` rather than raising when the directory
holds none, and it does not look one level down. The checkpointer writes the
export under `&lt;checkpoint&gt;/model/consolidated`, so accepting the checkpoint
root is what makes the config path usable, and a missing export has to fail
loudly: silently loading nothing would train stage 2 from a random draft.

**Parameters:**

**`draft_init_from`** `str`

A `.safetensors` file, or a directory holding the
consolidated export, or a directory one level above it.

---

**Returns:** `str`

The path to hand to `load_hf_safetensors_state_dict`.

**Raises:**

* `FileNotFoundError`: If no `.safetensors` file is at or under the path.

```python
nemo_automodel.recipes.llm.train_vispec._resolve_image_token_id(
    target_config,
    processor
) -> int
```

Resolve the target's image-placeholder token id.

**Parameters:**

**`target_config`**

The target VLM's HuggingFace config.

---

**`processor`**

The target's `AutoProcessor`.

---

**Returns:** `int`

The token id that marks an image position in `input_ids`.

```python
nemo_automodel.recipes.llm.train_vispec._seed_draft_initialization(
    recipe_cfg
) -> int
```

Seed ViSpec draft initialization and training-time stochasticity.

`shuffle_seed` predates the explicit `seed` knob and remains the
backwards-compatible default. Seeding must happen before the draft is
constructed because the frozen target is loaded from checkpoints while the
draft starts from random weights.

```python
nemo_automodel.recipes.llm.train_vispec.main(
    config_path: str | None = None
)
```

Entrypoint for the ViSpec stage selected by the recipe config.

**Parameters:**

**`config_path`** `str | None` — default: None

Optional default YAML path. The command-line `--config`
argument overrides this value.

---

**Raises:**

* `ValueError`: If `recipe` is not a supported ViSpec recipe name.

```python
nemo_automodel.recipes.llm.train_vispec.logger = logging.getLogger(__name__)
```