> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.vlm.finetune

## Module Contents

### Classes

| Name                                                                                                        | Description                                                               |
| ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- |
| [`FinetuneRecipeForVLM`](#nemo_automodel-recipes-vlm-finetune-FinetuneRecipeForVLM)                         | Recipe for fine-tuning a VLM model.                                       |
| [`_CpPackingCapability`](#nemo_automodel-recipes-vlm-finetune-_CpPackingCapability)                         | Model capability required by packed VLM context parallelism.              |
| [`_CpVisionFrameShardingCapability`](#nemo_automodel-recipes-vlm-finetune-_CpVisionFrameShardingCapability) | Model capability required by the VLM vision frame-sharding recipe policy. |

### Functions

| Name                                                                                                                            | Description                                                                        |
| ------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
| [`_accepted_targets`](#nemo_automodel-recipes-vlm-finetune-_accepted_targets)                                                   | Return the set of model `_target_` callables this recipe accepts.                  |
| [`_get_model_name`](#nemo_automodel-recipes-vlm-finetune-_get_model_name)                                                       | -                                                                                  |
| [`_is_recipe_target`](#nemo_automodel-recipes-vlm-finetune-_is_recipe_target)                                                   | True if `target` is on this recipe's allowlist of model entrypoints.               |
| [`_move_to_device`](#nemo_automodel-recipes-vlm-finetune-_move_to_device)                                                       | -                                                                                  |
| [`_shift_labels_left`](#nemo_automodel-recipes-vlm-finetune-_shift_labels_left)                                                 | Shift `labels` left by `k` positions, padding the tail with `-100`.                |
| [`_validate_cp_packing_support`](#nemo_automodel-recipes-vlm-finetune-_validate_cp_packing_support)                             | Reject packed CP before dataloader construction when routing is unsupported.       |
| [`_validate_cp_vision_frame_sharding_support`](#nemo_automodel-recipes-vlm-finetune-_validate_cp_vision_frame_sharding_support) | Reject enabled vision frame sharding when the model has no production integration. |
| [`build_dataloader`](#nemo_automodel-recipes-vlm-finetune-build_dataloader)                                                     | Build a DataLoader for the VLM dataset.                                            |
| [`build_model`](#nemo_automodel-recipes-vlm-finetune-build_model)                                                               | Build and initialize a model for VLM.                                              |
| [`main`](#nemo_automodel-recipes-vlm-finetune-main)                                                                             | Main entry point for the fine-tuning recipe.                                       |

### Data

[`logger`](#nemo_automodel-recipes-vlm-finetune-logger)

### API

```python
class nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM(
    cfg
)
```

**Bases:** [BaseRecipe](/nemo-automodel/nemo_automodel/recipes/base_recipe#nemo_automodel-recipes-base_recipe-BaseRecipe)

Recipe for fine-tuning a VLM model.

**`cfg`**

---

**`magi`** `= MagiState()`

---

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._configure_packing() -> nemo_automodel.components.models.common.packing.PackingCapabilities
```

Configure local model stages before the VLM dataloader is built.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._configure_pipeline_loss_fn()
```

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._cp_vision_frame_sharding_context()
```

Publish the CP-only group while a VLM forward may run its vision tower.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._create_distributed_setup() -> nemo_automodel.components.distributed.config.DistributedSetup
```

Create the distributed setup used by this recipe rank.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._forward_backward_step(
    idx,
    batch,
    loss_buffer,
    num_label_tokens,
    num_batches,
    is_train: bool = True
)
```

Run one local batch and accumulate its loss and optional gradients.

**Parameters:**

**`idx`**

Microbatch index in the accumulation window.

---

**`batch`**

Input mapping with token IDs, labels, and physical NEAT
document IDs of shape \[batch, sequence]. NEAT attention metadata
is batch-major; legacy THD inputs are flattened by the sharder.
VLM media and position tensors retain the model's input layout.
THD MTP requires physical boundaries in cu\_seqlens\_padded of
shape \[num\_sequences + 1] or \[1, num\_sequences + 1], or
\_packed\_seq\_ids \[batch, sequence] supplied by the native model
sharder from physical seq\_lens\_padded \[batch, num\_sequences].

---

**`loss_buffer`**

List receiving the detached scalar loss.

---

**`num_label_tokens`**

Global supervised-token count for loss normalization.

---

**`num_batches`**

Number of microbatches in the accumulation window.

---

**`is_train`** `bool` — default: True

Whether to backpropagate the combined main and MTP loss.

---

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._maybe_add_drafter_loss(
    out: typing.Any,
    base_loss: torch.Tensor,
    labels: torch.Tensor,
    model: torch.nn.Module,
    num_label_tokens: int,
    log: bool = False
) -> torch.Tensor
```

Return `base_loss + lambda * sum_k CE(drafter_logits[k], shifted_labels_k)`.

If `out` does not carry a non-empty `drafter_logits` attribute (i.e. the
model isn't a joint composite), returns `base_loss` unchanged.

For drafter step `k`, labels are shifted left by `k` positions to match
the VLM collate's pre-shifted convention (`labels[t] == input_ids[t+1]`).
`log=True` emits a one-line breakdown on rank 0; callers should gate this
on the appropriate step / microbatch index to avoid log spam.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._maybe_set_pp_first_stage_embed_input_meta(
    model_input: torch.Tensor
) -> None
```

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._run_train_optim_step(
    batches,
    max_grad_norm: float | None = None
)
```

Execute a single training step.

**Parameters:**

**`batches`**

List of batches of training data.

---

**`max_grad_norm`** `float | None` — default: None

Gradient clipping norm. Optional, if None will not clip gradients.

---

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._run_validation_epoch(
    val_dataloader
)
```

Run one pass over `self.val_dataloader`.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._should_setup_training_components() -> bool
```

Whether this rank owns the trainable model and its components.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.log_train_metrics(
    log_data
) -> float
```

Log metrics to wandb.

**Parameters:**

**`train_loss`**

Training loss.

---

**`grad_norm`**

Grad norm from the training step.

---

**`num_tokens_in_batch`**

Total number of loss tokens.

---

**`tps`**

Tokens per second.

---

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.log_val_metrics(
    log_data
)
```

Log metrics to wandb and other loggers
Args:
log\_data: MetricsSample object, containing:
step: int, the current step.
epoch: int, the current epoch.
metrics: Dict\[str, float], containing:
"val\_loss": Validation loss.
"lr": Learning rate.
"num\_label\_tokens": Number of label tokens.
"mem": Memory allocated.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.run_train_validation_loop()
```

Run the training loop over all epochs and batches.

For each batch, perform a forward pass, compute loss, backpropagate,
and update model parameters when necessary. Also prints loss every gradient step.

```python
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.setup()
```

Builds all components needed for training/validation/logging/checkpointing/etc.

This is the last place where self.cfg should be referenced.

**Raises:**

* `NotImplemented`: Raises if it tries to restore a checkpoint; will be removed.

```python
class nemo_automodel.recipes.vlm.finetune._CpPackingCapability()
```

Protocol

Model capability required by packed VLM context parallelism.

**`supports_cp_with_sequence_packing`** `bool`

Whether the model's active backend owns packed CP routing.

---

```python
class nemo_automodel.recipes.vlm.finetune._CpVisionFrameShardingCapability()
```

Protocol

Model capability required by the VLM vision frame-sharding recipe policy.

**`supports_cp_vision_frame_sharding`** `bool`

Whether the model owns a verified CP vision frame-sharding integration.

---

```python
nemo_automodel.recipes.vlm.finetune._accepted_targets() -> set
```

Return the set of model `_target_` callables this recipe accepts.

These are the wrapper-layer entrypoints that know how to absorb the
recipe's infrastructure kwargs (`device_mesh`, `distributed_config`,
`peft_config`, `freeze_config`, `pipeline_config`, plus the
optional `moe_config` / `fp8_config` / `compile_config`). Anything
not on this list is rejected with a clear error -- vanilla
`transformers.AutoModelFor*` does not handle these kwargs and would
otherwise fail deep inside HF code.

New infra-aware composites (e.g. Gemma4WithDrafter) opt in by adding their `.from_pretrained`
(and `.from_config` if applicable) here.

The Gemma4 joint composite is added behind a try/except because it
requires the optional `transformers.models.gemma4_assistant` module
that ships with `transformers&gt;=5.8.0.dev`.

```python
nemo_automodel.recipes.vlm.finetune._get_model_name(
    cfg_model
)
```

```python
nemo_automodel.recipes.vlm.finetune._is_recipe_target(
    target
) -> bool
```

True if `target` is on this recipe's allowlist of model entrypoints.

```python
nemo_automodel.recipes.vlm.finetune._move_to_device(
    value: typing.Any,
    device: torch.device
) -> typing.Any
```

```python
nemo_automodel.recipes.vlm.finetune._shift_labels_left(
    labels: torch.Tensor,
    k: int
) -> torch.Tensor
```

Shift `labels` left by `k` positions, padding the tail with `-100`.

Used to build drafter-step targets in joint base + drafter training.

The VLM collate pipeline already pre-shifts labels by 1 so that
`labels[t] == input_ids[t + 1]` (the next-token target). Drafter step `k`
predicts position `t + 1 + k` of the original sequence, which corresponds
to `labels[t + k]` in the pre-shifted convention. So for step `k`:

* `k = 0` (one-step drafter) -> no shift; reuse `labels` as-is.
* `k = 1` -> shift labels left by 1 (drafter predicts two tokens ahead).
* `k = n` -> shift labels left by `n`.

**Parameters:**

**`labels`** `torch.Tensor`

`[B, S]` LongTensor of label ids (`-100` marks ignored
positions).

---

**`k`** `int`

Number of positions to shift to the left. `k &lt;= 0` is a no-op.

---

**Returns:** `torch.Tensor`

A new `[B, S]` LongTensor with `labels[:, k:]` in the leading slice

```python
nemo_automodel.recipes.vlm.finetune._validate_cp_packing_support(
    model: nemo_automodel.recipes.vlm.finetune._CpPackingCapability,
    packing_enabled: bool,
    cp_size: int
) -> None
```

Reject packed CP before dataloader construction when routing is unsupported.

```python
nemo_automodel.recipes.vlm.finetune._validate_cp_vision_frame_sharding_support(
    model: nemo_automodel.recipes.vlm.finetune._CpVisionFrameShardingCapability,
    config: nemo_automodel.components.distributed.cp_vision_frame_shard.CpVisionFrameShardingConfig
) -> None
```

Reject enabled vision frame sharding when the model has no production integration.

```python
nemo_automodel.recipes.vlm.finetune.build_dataloader(
    cfg_ds,
    cfg_dl,
    pretrained_model_name_or_path,
    cfg_processor,
    device_mesh,
    seed,
    local_batch_size,
    cfg_model = None,
    cfg_ps = None,
    get_rope_index = None,
    pp_n_microbatches = None,
    model: torch.nn.Module | None = None
) -> tuple[torch.utils.data.DataLoader, transformers.processing_utils.ProcessorMixin]
```

Build a DataLoader for the VLM dataset.

**Parameters:**

**`cfg_ds`**

Dataset configuration.

---

**`cfg_dl`**

DataLoader configuration.

---

**`pretrained_model_name_or_path`**

Pretrained model name or path for processor loading.

---

**`cfg_processor`**

Processor configuration or None.

---

**`device_mesh`**

Device mesh for distributed training.

---

**`seed`**

Random seed.

---

**`local_batch_size`**

Local batch size.

---

**`cfg_model`** — default: None

Deprecated compatibility argument; ignored.

---

**`cfg_ps`** — default: None

Packed sequence configuration (top-level `packed_sequence:` section).
When provided, takes precedence over `dataset.packing`.

---

**`get_rope_index`** — default: None

Optional `model.get_rope_index` callable. When provided,
VLM neat packing computes mRoPE 3D position IDs per sample so packed
mRoPE-aware models (Qwen2.5-VL, Qwen3-VL, ...) preserve multimodal
position semantics across pack boundaries instead of falling back to
plain 1D positions.

---

**`pp_n_microbatches`** — default: None

When set, wrap collate so VLM media tensors are
pre-chunked for this many PP microbatches before entering the train loop.

---

**`model`** `nn.Module | None` — default: None

Built model supplying the structural packing contract.

---

**Returns:** `tuple[DataLoader, ProcessorMixin]`

The instantiated DataLoader and processor.

```python
nemo_automodel.recipes.vlm.finetune.build_model(
    cfg_model,
    cfg_freeze,
    cfg_peft,
    seed,
    cfg_fp8 = None,
    cfg_compile = None,
    distributed_setup: nemo_automodel.components.distributed.config.DistributedSetup | None = None,
    cfg_quantization = None
) -> tuple[torch.nn.Module | nemo_automodel.components.distributed.pipelining.AutoPipeline, list['Optimizer']]
```

Build and initialize a model for VLM.

**Returns:** `tuple[nn.Module | AutoPipeline, list['Optimizer']]`

The instantiated model and optimizer.

```python
nemo_automodel.recipes.vlm.finetune.main(
    config_path = None
)
```

Main entry point for the fine-tuning recipe.

Loads the configuration, sets up the trainer, and initiates the training loop.

```python
nemo_automodel.recipes.vlm.finetune.logger = logging.getLogger(__name__)
```