nemo_automodel.components.models.minimax_m3_vl.flops

View as Markdown

Useful text-backbone FLOPs for MiniMax M3 training.

Module Contents

Functions

NameDescription
model_flopsCount useful GEMM FLOPs for text-only full-parameter training.

API

nemo_automodel.components.models.minimax_m3_vl.flops.model_flops(
gbs: int = 1,
seq_len: int | None = None
) -> float

Count useful GEMM FLOPs for text-only full-parameter training.

Trainable projections and attention count forward, input gradients and weight gradients (six FLOPs per MAC). Hard block selection has no gradient, so the indexer projections and causal scores count forward only (two per MAC). Activation recomputation, masked-out dense work, padding, optimizer updates, communication, softmax and elementwise operations are excluded. This is MFU, not the FLOPs actually executed by a particular backend (HFU).

Each sequence is one document without padding. Sparse layers force the current block into the top-k budget; its causal partial length is counted exactly. Packed-document FLOPs require the individual document lengths and cannot be inferred from a packed row length.

Parameters:

config
MiniMaxM3VLTextConfig

Effective text-backbone configuration, with MTP disabled.

gbs
intDefaults to 1

Number of full-length sequences in the global optimizer step.

seq_len
int | NoneDefaults to None

Tokens in each sequence; defaults to max_position_embeddings.

Returns: float

Useful model FLOPs over all devices for one optimizer step.

Raises:

  • ValueError: Sequence dimensions, MTP, or sparse selection settings cannot be represented by this text-backbone accounting.