> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.multimodal.finetune

Fine-tuning recipe for the BAGEL multimodal family.

BAGEL uses packed mixed-modality batches and returns `dict(ce=..., mse=...)`,
so this recipe subclasses `BaseRecipe` directly instead of the standard
VLM recipe. It supports Stage 1 understanding-only CE and Stage 2 joint
understanding + visual generation with VAE encode and flow-matching MSE.

Key training-step behavior:

* Per-token CE is reduced via `ce.sum() * world_size / total_ce_tokens`.
* Per-token MSE is reduced via
  `mse.mean(dim=-1).sum() * world_size / total_mse_tokens`.
* Optimizer: AdamW(lr=2e-5, betas=(0.9, 0.95), eps=1e-15, weight\_decay=0).
* LR schedule: constant with 2000-step warmup.
* bf16 autocast on the forward pass; FSDP2/HSDP via AM's model
  infrastructure using the `distributed.strategy: fsdp2` YAML knob.
* Global seed = 4396, data seed = 42 (BAGEL defaults).

## Module Contents

### Classes

| Name                                                                                                     | Description                                                   |
| -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
| [`FinetuneRecipeForMultimodal`](#nemo_automodel-recipes-multimodal-finetune-FinetuneRecipeForMultimodal) | Fine-tuning recipe for BAGEL packed Stage 1/Stage 2 training. |

### Functions

| Name                                                                                                         | Description                                                               |
| ------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- |
| [`_first_path`](#nemo_automodel-recipes-multimodal-finetune-_first_path)                                     | Return the first non-empty string-like path from `values`.                |
| [`_load_bagel_tokenizer`](#nemo_automodel-recipes-multimodal-finetune-_load_bagel_tokenizer)                 | Load the BAGEL Qwen2 tokenizer and add the four special tokens.           |
| [`_load_bagel_vae`](#nemo_automodel-recipes-multimodal-finetune-_load_bagel_vae)                             | Load BAGEL's autoencoder from AM-owned code.                              |
| [`_maybe_resize_bagel_vocab`](#nemo_automodel-recipes-multimodal-finetune-_maybe_resize_bagel_vocab)         | Resize BAGEL token embeddings only when the tokenizer is larger.          |
| [`_resolve_bagel_artifact_path`](#nemo_automodel-recipes-multimodal-finetune-_resolve_bagel_artifact_path)   | Resolve the BAGEL artifact source used for config/tokenizer/VAE sidecars. |
| [`_resolve_bagel_tokenizer_path`](#nemo_automodel-recipes-multimodal-finetune-_resolve_bagel_tokenizer_path) | Resolve the tokenizer source for BAGEL training.                          |
| [`_resolve_bagel_vae_path`](#nemo_automodel-recipes-multimodal-finetune-_resolve_bagel_vae_path)             | Resolve `ae.safetensors` from a local checkpoint directory or HF repo.    |
| [`main`](#nemo_automodel-recipes-multimodal-finetune-main)                                                   | Run the BAGEL multimodal training recipe from a YAML config path.         |

### Data

[`_BAGEL_STAGE1_FORWARD_KEYS`](#nemo_automodel-recipes-multimodal-finetune-_BAGEL_STAGE1_FORWARD_KEYS)

[`_BAGEL_STAGE2_FORWARD_KEYS`](#nemo_automodel-recipes-multimodal-finetune-_BAGEL_STAGE2_FORWARD_KEYS)

[`logger`](#nemo_automodel-recipes-multimodal-finetune-logger)

### API

```python
class nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal(
    cfg
)
```

**Bases:** [BaseRecipe](/nemo-automodel/nemo_automodel/recipes/base_recipe#nemo_automodel-recipes-base_recipe-BaseRecipe)

Fine-tuning recipe for BAGEL packed Stage 1/Stage 2 training.

Subclasses `BaseRecipe` directly (rather than
`FinetuneRecipeForVLM`)
because BAGEL's packed-sequence input schema and
`dict(ce=..., mse=...)` forward-return are incompatible with the
standard VLM training loop. The VLM recipe was evaluated as a parent
class; the override surface ended up at \~95% of the class body, so a
parallel implementation is cleaner.

**`_last_tokens_per_sec`** `= 0.0`

---

**`_last_tokens_per_step`** `= 0.0`

---

**`_last_train_steps_per_sec`** `= 0.0`

---

**`_throughput_step_window`** `= 0`

---

**`_throughput_token_window`** `= 0.0`

---

**`_throughput_window_start`** `float | None = None`

---

**`_vae_path`** `str | None = None`

---

**`vae_encode_micro_batch_size`**

---

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._apply_warmup(
    step: int
) -> None
```

BAGEL uses constant LR with linear warmup from 0 to base\_lr.
Matches HF get\_constant\_schedule\_with\_warmup: scale = step/warmup\_steps
(so the first optimizer step at step=0 runs with lr=0).

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._build_hf_backbone_bagel_model(
    artifact_path: str | None,
    stage: int,
    rank_seed: int,
    init_seed: int,
    freeze_before_infrastructure: bool = False
)
```

Build BAGEL from HF backbones and apply the configured infrastructure.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._build_wandb()
```

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._copy_vae_sidecar_to_checkpoint(
    checkpoint_path: pathlib.Path
) -> None
```

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._encode_vae_images(
    padded_images: torch.Tensor
) -> torch.Tensor
```

Encode VAE images in chunks to cap frozen-VAE activation peaks.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._forward_backward_step(
    idx: int,
    batch,
    loss_buffer: typing.List[torch.Tensor],
    num_ce_tokens_global: int,
    num_mse_tokens_global: int = 0,
    num_batches: int,
    is_train: bool = True
)
```

One packed-sample forward + backward.

Replaces the parent VLM step because BAGEL's forward takes a packed-
sequence kwarg dict and returns `&#123;"ce": Tensor|None, "mse": Tensor|
None&#125;` instead of the HF `ModelOutput` the VLM recipe expects.

Stage 2: VAE-encodes `padded_images` into `padded_latent` before
the model forward, then composes `ce + mse` reductions into one
microbatch loss.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._prepare_batch(
    batch
) -> typing.Dict[str, typing.Any]
```

Move the SimpleCustomBatch packed dict onto the current CUDA device.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._run_train_optim_step(
    batches,
    max_grad_norm: float | None = None
)
```

Execute a training step; supports grad accumulation trivially.

BAGEL packs variable token counts per microbatch. For gradient
accumulation, normalize each CE/MSE contribution by the token count
over the whole optimizer step, not by each microbatch independently.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._use_sharded_model_ema(
    ema_impl: str,
    model_init_mode: str
) -> bool
```

Resolve the EMA implementation choice.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.log_train_metrics(
    log_data
) -> None
```

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.run_train_validation_loop()
```

BAGEL training loop — no validation, optional periodic checkpoint.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.save_checkpoint(
    epoch: int,
    step: int,
    train_loss: float,
    val_loss: typing.Dict[str, float] | None = None,
    best_metric_key: str = 'default'
) -> None
```

Save BAGEL state and include the frozen VAE as a checkpoint sidecar.

```python
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.setup()
```

Build distributed, model, tokenizer, data, optimizer, scheduler.

```python
nemo_automodel.recipes.multimodal.finetune._first_path(
    values: typing.Any = ()
) -> str | None
```

Return the first non-empty string-like path from `values`.

```python
nemo_automodel.recipes.multimodal.finetune._load_bagel_tokenizer(
    model_path: str
)
```

Load the BAGEL Qwen2 tokenizer and add the four special tokens.

Returns `(tokenizer, special_tokens_dict, num_new_tokens)`.

```python
nemo_automodel.recipes.multimodal.finetune._load_bagel_vae(
    vae_path: str
) -> tuple[typing.Any, typing.Any]
```

Load BAGEL's autoencoder from AM-owned code.

```python
nemo_automodel.recipes.multimodal.finetune._maybe_resize_bagel_vocab(
    model,
    tokenizer_vocab_size: int,
    num_new_tokens: int
) -> None
```

Resize BAGEL token embeddings only when the tokenizer is larger.

The released BAGEL checkpoint intentionally has a padded vocab
(`152064`) larger than the tokenizer length. Shrinking to the tokenizer
length removes trainable rows from both `embed_tokens` and `lm_head`,
which changes the trainable parameter set.

```python
nemo_automodel.recipes.multimodal.finetune._resolve_bagel_artifact_path(
    cfg
) -> str | None
```

Resolve the BAGEL artifact source used for config/tokenizer/VAE sidecars.

Fine-tuning usually uses `model.pretrained_model_name_or_path`. Pretraining
can instead use `model.config.pretrained_model_name_or_path` with
`NeMoAutoModelForMultimodalLM.from_config` so no base weights are loaded,
while still resolving tokenizer/config/VAE side artifacts.

```python
nemo_automodel.recipes.multimodal.finetune._resolve_bagel_tokenizer_path(
    cfg
) -> str
```

Resolve the tokenizer source for BAGEL training.

```python
nemo_automodel.recipes.multimodal.finetune._resolve_bagel_vae_path(
    model_path: str | None,
    vae_path: str | None
) -> str
```

Resolve `ae.safetensors` from a local checkpoint directory or HF repo.

```python
nemo_automodel.recipes.multimodal.finetune.main(
    config_path: str | None = None
) -> None
```

Run the BAGEL multimodal training recipe from a YAML config path.

```python
nemo_automodel.recipes.multimodal.finetune._BAGEL_STAGE1_FORWARD_KEYS = {'sequence_length', 'packed_text_ids', 'packed_text_indexes', 'sample_lens', 'pa...
```

```python
nemo_automodel.recipes.multimodal.finetune._BAGEL_STAGE2_FORWARD_KEYS = _BAGEL_STAGE1_FORWARD_KEYS | {'padded_latent', 'patchified_vae_latent_shapes', '...
```

```python
nemo_automodel.recipes.multimodal.finetune.logger = logging.getLogger(__name__)
```