nemo_automodel.components.models.minimax_m3_vl.flops
nemo_automodel.components.models.minimax_m3_vl.flops
Useful text-backbone FLOPs for MiniMax M3 training.
Module Contents
Functions
API
Count useful GEMM FLOPs for text-only full-parameter training.
Trainable projections and attention count forward, input gradients and weight gradients (six FLOPs per MAC). Hard block selection has no gradient, so the indexer projections and causal scores count forward only (two per MAC). Activation recomputation, masked-out dense work, padding, optimizer updates, communication, softmax and elementwise operations are excluded. This is MFU, not the FLOPs actually executed by a particular backend (HFU).
Each sequence is one document without padding. Sparse layers force the current block into the top-k budget; its causal partial length is counted exactly. Packed-document FLOPs require the individual document lengths and cannot be inferred from a packed row length.
Parameters:
Effective text-backbone configuration, with MTP disabled.
Number of full-length sequences in the global optimizer step.
Tokens in each sequence; defaults to max_position_embeddings.
Returns: float
Useful model FLOPs over all devices for one optimizer step.
Raises:
ValueError: Sequence dimensions, MTP, or sparse selection settings cannot be represented by this text-backbone accounting.