nemo_automodel.recipes.vlm.finetune

View as Markdown

Module Contents

Classes

NameDescription
FinetuneRecipeForVLMRecipe for fine-tuning a VLM model.
_CpPackingCapabilityModel capability required by packed VLM context parallelism.
_CpVisionFrameShardingCapabilityModel capability required by the VLM vision frame-sharding recipe policy.

Functions

NameDescription
_accepted_targetsReturn the set of model _target_ callables this recipe accepts.
_get_model_name-
_is_recipe_targetTrue if target is on this recipe’s allowlist of model entrypoints.
_move_to_device-
_shift_labels_leftShift labels left by k positions, padding the tail with -100.
_validate_cp_packing_supportReject packed CP before dataloader construction when routing is unsupported.
_validate_cp_vision_frame_sharding_supportReject enabled vision frame sharding when the model has no production integration.
build_dataloaderBuild a DataLoader for the VLM dataset.
build_modelBuild and initialize a model for VLM.
mainMain entry point for the fine-tuning recipe.

Data

logger

API

class nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM(
cfg
)

Bases: BaseRecipe

Recipe for fine-tuning a VLM model.

cfg
magi
= MagiState()
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._configure_packing() -> nemo_automodel.components.models.common.packing.PackingCapabilities

Configure local model stages before the VLM dataloader is built.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._configure_pipeline_loss_fn()
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._cp_vision_frame_sharding_context()

Publish the CP-only group while a VLM forward may run its vision tower.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._create_distributed_setup() -> nemo_automodel.components.distributed.config.DistributedSetup

Create the distributed setup used by this recipe rank.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._forward_backward_step(
idx,
batch,
loss_buffer,
num_label_tokens,
num_batches,
is_train: bool = True
)

Run one local batch and accumulate its loss and optional gradients.

Parameters:

idx

Microbatch index in the accumulation window.

batch

Input mapping with token IDs, labels, and physical NEAT document IDs of shape [batch, sequence]. NEAT attention metadata is batch-major; legacy THD inputs are flattened by the sharder. VLM media and position tensors retain the model’s input layout. THD MTP requires physical boundaries in cu_seqlens_padded of shape [num_sequences + 1] or [1, num_sequences + 1], or _packed_seq_ids [batch, sequence] supplied by the native model sharder from physical seq_lens_padded [batch, num_sequences].

loss_buffer

List receiving the detached scalar loss.

num_label_tokens

Global supervised-token count for loss normalization.

num_batches

Number of microbatches in the accumulation window.

is_train
boolDefaults to True

Whether to backpropagate the combined main and MTP loss.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._maybe_add_drafter_loss(
out: typing.Any,
base_loss: torch.Tensor,
labels: torch.Tensor,
model: torch.nn.Module,
num_label_tokens: int,
log: bool = False
) -> torch.Tensor

Return base_loss + lambda * sum_k CE(drafter_logits[k], shifted_labels_k).

If out does not carry a non-empty drafter_logits attribute (i.e. the model isn’t a joint composite), returns base_loss unchanged.

For drafter step k, labels are shifted left by k positions to match the VLM collate’s pre-shifted convention (labels[t] == input_ids[t+1]). log=True emits a one-line breakdown on rank 0; callers should gate this on the appropriate step / microbatch index to avoid log spam.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._maybe_set_pp_first_stage_embed_input_meta(
model_input: torch.Tensor
) -> None
nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._run_train_optim_step(
batches,
max_grad_norm: float | None = None
)

Execute a single training step.

Parameters:

batches

List of batches of training data.

max_grad_norm
float | NoneDefaults to None

Gradient clipping norm. Optional, if None will not clip gradients.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._run_validation_epoch(
val_dataloader
)

Run one pass over self.val_dataloader.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM._should_setup_training_components() -> bool

Whether this rank owns the trainable model and its components.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.log_train_metrics(
log_data
) -> float

Log metrics to wandb.

Parameters:

train_loss

Training loss.

grad_norm

Grad norm from the training step.

num_tokens_in_batch

Total number of loss tokens.

tps

Tokens per second.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.log_val_metrics(
log_data
)

Log metrics to wandb and other loggers Args: log_data: MetricsSample object, containing: step: int, the current step. epoch: int, the current epoch. metrics: Dict[str, float], containing: “val_loss”: Validation loss. “lr”: Learning rate. “num_label_tokens”: Number of label tokens. “mem”: Memory allocated.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.run_train_validation_loop()

Run the training loop over all epochs and batches.

For each batch, perform a forward pass, compute loss, backpropagate, and update model parameters when necessary. Also prints loss every gradient step.

nemo_automodel.recipes.vlm.finetune.FinetuneRecipeForVLM.setup()

Builds all components needed for training/validation/logging/checkpointing/etc.

This is the last place where self.cfg should be referenced.

Raises:

  • NotImplemented: Raises if it tries to restore a checkpoint; will be removed.
class nemo_automodel.recipes.vlm.finetune._CpPackingCapability()
Protocol

Model capability required by packed VLM context parallelism.

supports_cp_with_sequence_packing
bool

Whether the model’s active backend owns packed CP routing.

class nemo_automodel.recipes.vlm.finetune._CpVisionFrameShardingCapability()
Protocol

Model capability required by the VLM vision frame-sharding recipe policy.

supports_cp_vision_frame_sharding
bool

Whether the model owns a verified CP vision frame-sharding integration.

nemo_automodel.recipes.vlm.finetune._accepted_targets() -> set

Return the set of model _target_ callables this recipe accepts.

These are the wrapper-layer entrypoints that know how to absorb the recipe’s infrastructure kwargs (device_mesh, distributed_config, peft_config, freeze_config, pipeline_config, plus the optional moe_config / fp8_config / compile_config). Anything not on this list is rejected with a clear error — vanilla transformers.AutoModelFor* does not handle these kwargs and would otherwise fail deep inside HF code.

New infra-aware composites (e.g. Gemma4WithDrafter) opt in by adding their .from_pretrained (and .from_config if applicable) here.

The Gemma4 joint composite is added behind a try/except because it requires the optional transformers.models.gemma4_assistant module that ships with transformers>=5.8.0.dev.

nemo_automodel.recipes.vlm.finetune._get_model_name(
cfg_model
)
nemo_automodel.recipes.vlm.finetune._is_recipe_target(
target
) -> bool

True if target is on this recipe’s allowlist of model entrypoints.

nemo_automodel.recipes.vlm.finetune._move_to_device(
value: typing.Any,
device: torch.device
) -> typing.Any
nemo_automodel.recipes.vlm.finetune._shift_labels_left(
labels: torch.Tensor,
k: int
) -> torch.Tensor

Shift labels left by k positions, padding the tail with -100.

Used to build drafter-step targets in joint base + drafter training.

The VLM collate pipeline already pre-shifts labels by 1 so that labels[t] == input_ids[t + 1] (the next-token target). Drafter step k predicts position t + 1 + k of the original sequence, which corresponds to labels[t + k] in the pre-shifted convention. So for step k:

  • k = 0 (one-step drafter) -> no shift; reuse labels as-is.
  • k = 1 -> shift labels left by 1 (drafter predicts two tokens ahead).
  • k = n -> shift labels left by n.

Parameters:

labels
torch.Tensor

[B, S] LongTensor of label ids (-100 marks ignored positions).

k
int

Number of positions to shift to the left. k <= 0 is a no-op.

Returns: torch.Tensor

A new [B, S] LongTensor with labels[:, k:] in the leading slice

nemo_automodel.recipes.vlm.finetune._validate_cp_packing_support(
packing_enabled: bool,
cp_size: int
) -> None

Reject packed CP before dataloader construction when routing is unsupported.

nemo_automodel.recipes.vlm.finetune._validate_cp_vision_frame_sharding_support(
) -> None

Reject enabled vision frame sharding when the model has no production integration.

nemo_automodel.recipes.vlm.finetune.build_dataloader(
cfg_ds,
cfg_dl,
pretrained_model_name_or_path,
cfg_processor,
device_mesh,
seed,
local_batch_size,
cfg_model = None,
cfg_ps = None,
get_rope_index = None,
pp_n_microbatches = None,
model: torch.nn.Module | None = None
) -> tuple[torch.utils.data.DataLoader, transformers.processing_utils.ProcessorMixin]

Build a DataLoader for the VLM dataset.

Parameters:

cfg_ds

Dataset configuration.

cfg_dl

DataLoader configuration.

pretrained_model_name_or_path

Pretrained model name or path for processor loading.

cfg_processor

Processor configuration or None.

device_mesh

Device mesh for distributed training.

seed

Random seed.

local_batch_size

Local batch size.

cfg_model
Defaults to None

Deprecated compatibility argument; ignored.

cfg_ps
Defaults to None

Packed sequence configuration (top-level packed_sequence: section). When provided, takes precedence over dataset.packing.

get_rope_index
Defaults to None

Optional model.get_rope_index callable. When provided, VLM neat packing computes mRoPE 3D position IDs per sample so packed mRoPE-aware models (Qwen2.5-VL, Qwen3-VL, …) preserve multimodal position semantics across pack boundaries instead of falling back to plain 1D positions.

pp_n_microbatches
Defaults to None

When set, wrap collate so VLM media tensors are pre-chunked for this many PP microbatches before entering the train loop.

model
nn.Module | NoneDefaults to None

Built model supplying the structural packing contract.

Returns: tuple[DataLoader, ProcessorMixin]

The instantiated DataLoader and processor.

nemo_automodel.recipes.vlm.finetune.build_model(
cfg_model,
cfg_freeze,
cfg_peft,
seed,
cfg_fp8 = None,
cfg_compile = None,
cfg_quantization = None
) -> tuple[torch.nn.Module | nemo_automodel.components.distributed.pipelining.AutoPipeline, list['Optimizer']]

Build and initialize a model for VLM.

Returns: tuple[nn.Module | AutoPipeline, list['Optimizer']]

The instantiated model and optimizer.

nemo_automodel.recipes.vlm.finetune.main(
config_path = None
)

Main entry point for the fine-tuning recipe.

Loads the configuration, sets up the trainer, and initiates the training loop.

nemo_automodel.recipes.vlm.finetune.logger = logging.getLogger(__name__)