nemo_automodel.recipes.multimodal.finetune

View as Markdown

Fine-tuning recipe for the BAGEL multimodal family.

BAGEL uses packed mixed-modality batches and returns dict(ce=..., mse=...), so this recipe subclasses BaseRecipe directly instead of the standard VLM recipe. It supports Stage 1 understanding-only CE and Stage 2 joint understanding + visual generation with VAE encode and flow-matching MSE.

Key training-step behavior:

  • Per-token CE is reduced via ce.sum() * world_size / total_ce_tokens.
  • Per-token MSE is reduced via mse.mean(dim=-1).sum() * world_size / total_mse_tokens.
  • Optimizer: AdamW(lr=2e-5, betas=(0.9, 0.95), eps=1e-15, weight_decay=0).
  • LR schedule: constant with 2000-step warmup.
  • bf16 autocast on the forward pass; FSDP2/HSDP via AM’s model infrastructure using the distributed.strategy: fsdp2 YAML knob.
  • Global seed = 4396, data seed = 42 (BAGEL defaults).

Module Contents

Classes

NameDescription
FinetuneRecipeForMultimodalFine-tuning recipe for BAGEL packed Stage 1/Stage 2 training.

Functions

NameDescription
_first_pathReturn the first non-empty string-like path from values.
_load_bagel_tokenizerLoad the BAGEL Qwen2 tokenizer and add the four special tokens.
_load_bagel_vaeLoad BAGEL’s autoencoder from AM-owned code.
_maybe_resize_bagel_vocabResize BAGEL token embeddings only when the tokenizer is larger.
_resolve_bagel_artifact_pathResolve the BAGEL artifact source used for config/tokenizer/VAE sidecars.
_resolve_bagel_tokenizer_pathResolve the tokenizer source for BAGEL training.
_resolve_bagel_vae_pathResolve ae.safetensors from a local checkpoint directory or HF repo.
mainRun the BAGEL multimodal training recipe from a YAML config path.

Data

_BAGEL_STAGE1_FORWARD_KEYS

_BAGEL_STAGE2_FORWARD_KEYS

logger

API

class nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal(
cfg
)

Bases: BaseRecipe

Fine-tuning recipe for BAGEL packed Stage 1/Stage 2 training.

Subclasses BaseRecipe directly (rather than FinetuneRecipeForVLM) because BAGEL’s packed-sequence input schema and dict(ce=..., mse=...) forward-return are incompatible with the standard VLM training loop. The VLM recipe was evaluated as a parent class; the override surface ended up at ~95% of the class body, so a parallel implementation is cleaner.

_last_tokens_per_sec
= 0.0
_last_tokens_per_step
= 0.0
_last_train_steps_per_sec
= 0.0
_throughput_step_window
= 0
_throughput_token_window
= 0.0
_throughput_window_start
float | None = None
_vae_path
str | None = None
vae_encode_micro_batch_size
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._apply_warmup(
step: int
) -> None

BAGEL uses constant LR with linear warmup from 0 to base_lr. Matches HF get_constant_schedule_with_warmup: scale = step/warmup_steps (so the first optimizer step at step=0 runs with lr=0).

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._build_hf_backbone_bagel_model(
artifact_path: str | None,
stage: int,
rank_seed: int,
init_seed: int,
freeze_before_infrastructure: bool = False
)

Build BAGEL from HF backbones and apply the configured infrastructure.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._build_wandb()
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._copy_vae_sidecar_to_checkpoint(
checkpoint_path: pathlib.Path
) -> None
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._encode_vae_images(
padded_images: torch.Tensor
) -> torch.Tensor

Encode VAE images in chunks to cap frozen-VAE activation peaks.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._forward_backward_step(
idx: int,
batch,
loss_buffer: typing.List[torch.Tensor],
num_ce_tokens_global: int,
num_mse_tokens_global: int = 0,
num_batches: int,
is_train: bool = True
)

One packed-sample forward + backward.

Replaces the parent VLM step because BAGEL’s forward takes a packed- sequence kwarg dict and returns {"ce": Tensor|None, "mse": Tensor| None} instead of the HF ModelOutput the VLM recipe expects.

Stage 2: VAE-encodes padded_images into padded_latent before the model forward, then composes ce + mse reductions into one microbatch loss.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._prepare_batch(
batch
) -> typing.Dict[str, typing.Any]

Move the SimpleCustomBatch packed dict onto the current CUDA device.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._run_train_optim_step(
batches,
max_grad_norm: float | None = None
)

Execute a training step; supports grad accumulation trivially.

BAGEL packs variable token counts per microbatch. For gradient accumulation, normalize each CE/MSE contribution by the token count over the whole optimizer step, not by each microbatch independently.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal._use_sharded_model_ema(
ema_impl: str,
model_init_mode: str
) -> bool

Resolve the EMA implementation choice.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.log_train_metrics(
log_data
) -> None
nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.run_train_validation_loop()

BAGEL training loop — no validation, optional periodic checkpoint.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.save_checkpoint(
epoch: int,
step: int,
train_loss: float,
val_loss: typing.Dict[str, float] | None = None,
best_metric_key: str = 'default'
) -> None

Save BAGEL state and include the frozen VAE as a checkpoint sidecar.

nemo_automodel.recipes.multimodal.finetune.FinetuneRecipeForMultimodal.setup()

Build distributed, model, tokenizer, data, optimizer, scheduler.

nemo_automodel.recipes.multimodal.finetune._first_path(
values: typing.Any = ()
) -> str | None

Return the first non-empty string-like path from values.

nemo_automodel.recipes.multimodal.finetune._load_bagel_tokenizer(
model_path: str
)

Load the BAGEL Qwen2 tokenizer and add the four special tokens.

Returns (tokenizer, special_tokens_dict, num_new_tokens).

nemo_automodel.recipes.multimodal.finetune._load_bagel_vae(
vae_path: str
) -> tuple[typing.Any, typing.Any]

Load BAGEL’s autoencoder from AM-owned code.

nemo_automodel.recipes.multimodal.finetune._maybe_resize_bagel_vocab(
model,
tokenizer_vocab_size: int,
num_new_tokens: int
) -> None

Resize BAGEL token embeddings only when the tokenizer is larger.

The released BAGEL checkpoint intentionally has a padded vocab (152064) larger than the tokenizer length. Shrinking to the tokenizer length removes trainable rows from both embed_tokens and lm_head, which changes the trainable parameter set.

nemo_automodel.recipes.multimodal.finetune._resolve_bagel_artifact_path(
cfg
) -> str | None

Resolve the BAGEL artifact source used for config/tokenizer/VAE sidecars.

Fine-tuning usually uses model.pretrained_model_name_or_path. Pretraining can instead use model.config.pretrained_model_name_or_path with NeMoAutoModelForMultimodalLM.from_config so no base weights are loaded, while still resolving tokenizer/config/VAE side artifacts.

nemo_automodel.recipes.multimodal.finetune._resolve_bagel_tokenizer_path(
cfg
) -> str

Resolve the tokenizer source for BAGEL training.

nemo_automodel.recipes.multimodal.finetune._resolve_bagel_vae_path(
model_path: str | None,
vae_path: str | None
) -> str

Resolve ae.safetensors from a local checkpoint directory or HF repo.

nemo_automodel.recipes.multimodal.finetune.main(
config_path: str | None = None
) -> None

Run the BAGEL multimodal training recipe from a YAML config path.

nemo_automodel.recipes.multimodal.finetune._BAGEL_STAGE1_FORWARD_KEYS = {'sequence_length', 'packed_text_ids', 'packed_text_indexes', 'sample_lens', 'pa...
nemo_automodel.recipes.multimodal.finetune._BAGEL_STAGE2_FORWARD_KEYS = _BAGEL_STAGE1_FORWARD_KEYS | {'padded_latent', 'patchified_vae_latent_shapes', '...
nemo_automodel.recipes.multimodal.finetune.logger = logging.getLogger(__name__)