nemo_automodel.recipes.multimodal.finetune
nemo_automodel.recipes.multimodal.finetune
Fine-tuning recipe for the BAGEL multimodal family.
BAGEL uses packed mixed-modality batches and returns dict(ce=..., mse=...),
so this recipe subclasses BaseRecipe directly instead of the standard
VLM recipe. It supports Stage 1 understanding-only CE and Stage 2 joint
understanding + visual generation with VAE encode and flow-matching MSE.
Key training-step behavior:
- Per-token CE is reduced via
ce.sum() * world_size / total_ce_tokens. - Per-token MSE is reduced via
mse.mean(dim=-1).sum() * world_size / total_mse_tokens. - Optimizer: AdamW(lr=2e-5, betas=(0.9, 0.95), eps=1e-15, weight_decay=0).
- LR schedule: constant with 2000-step warmup.
- bf16 autocast on the forward pass; FSDP2/HSDP via AM’s model
infrastructure using the
distributed.strategy: fsdp2YAML knob. - Global seed = 4396, data seed = 42 (BAGEL defaults).
Module Contents
Classes
Functions
Data
API
Bases: BaseRecipe
Fine-tuning recipe for BAGEL packed Stage 1/Stage 2 training.
Subclasses BaseRecipe directly (rather than
FinetuneRecipeForVLM)
because BAGEL’s packed-sequence input schema and
dict(ce=..., mse=...) forward-return are incompatible with the
standard VLM training loop. The VLM recipe was evaluated as a parent
class; the override surface ended up at ~95% of the class body, so a
parallel implementation is cleaner.
BAGEL uses constant LR with linear warmup from 0 to base_lr. Matches HF get_constant_schedule_with_warmup: scale = step/warmup_steps (so the first optimizer step at step=0 runs with lr=0).
Build BAGEL from HF backbones and apply the configured infrastructure.
Encode VAE images in chunks to cap frozen-VAE activation peaks.
One packed-sample forward + backward.
Replaces the parent VLM step because BAGEL’s forward takes a packed-
sequence kwarg dict and returns {"ce": Tensor|None, "mse": Tensor| None} instead of the HF ModelOutput the VLM recipe expects.
Stage 2: VAE-encodes padded_images into padded_latent before
the model forward, then composes ce + mse reductions into one
microbatch loss.
Move the SimpleCustomBatch packed dict onto the current CUDA device.
Execute a training step; supports grad accumulation trivially.
BAGEL packs variable token counts per microbatch. For gradient accumulation, normalize each CE/MSE contribution by the token count over the whole optimizer step, not by each microbatch independently.
Resolve the EMA implementation choice.
BAGEL training loop — no validation, optional periodic checkpoint.
Save BAGEL state and include the frozen VAE as a checkpoint sidecar.
Build distributed, model, tokenizer, data, optimizer, scheduler.
Return the first non-empty string-like path from values.
Load the BAGEL Qwen2 tokenizer and add the four special tokens.
Returns (tokenizer, special_tokens_dict, num_new_tokens).
Load BAGEL’s autoencoder from AM-owned code.
Resize BAGEL token embeddings only when the tokenizer is larger.
The released BAGEL checkpoint intentionally has a padded vocab
(152064) larger than the tokenizer length. Shrinking to the tokenizer
length removes trainable rows from both embed_tokens and lm_head,
which changes the trainable parameter set.
Resolve the BAGEL artifact source used for config/tokenizer/VAE sidecars.
Fine-tuning usually uses model.pretrained_model_name_or_path. Pretraining
can instead use model.config.pretrained_model_name_or_path with
NeMoAutoModelForMultimodalLM.from_config so no base weights are loaded,
while still resolving tokenizer/config/VAE side artifacts.
Resolve the tokenizer source for BAGEL training.
Resolve ae.safetensors from a local checkpoint directory or HF repo.
Run the BAGEL multimodal training recipe from a YAML config path.