nemo_automodel.components.utils.flops_utils
nemo_automodel.components.utils.flops_utils
Module Contents
Functions
API
Build a list of 0/1 indicating dense(0) vs MoE(1) per layer.
Handles multiple config styles: first_k_dense_replace + moe_layer_freq, mlp_layer_types list, etc.
FLOPs for a single Gated DeltaNet (GDN / linear attention) layer.
Based on the GDN FLOPs calculator from Megatron-Bridge PR #2925.
Model FLOPs for hybrid model
Per-layer FLOPs for Kimi Delta Attention (KDA, gated delta-rule linear attention).
Projections follow the KDA layer definition (Kimi Linear, arXiv:2510.26692):
q/k/v with short causal convolutions, the low-rank forget gate
(f_a/f_b), the beta projection, the output gate (full-rank g_proj
or low-rank g_a/g_b) and o_proj. The chunkwise kernel is costed
with the paper’s own count, FLOPs_KDA(T; C, d_h) = 6 T d_h^2 + 3 T C d_h + T C^2
per head (forward, chunk size C), halved to MAC-equivalents so the shared
6 x (forward + backward) factor of this module applies.
Model FLOPs for Mamba layer.
Multiplied by 6 (3x fwd+bwd * 2x FMA) for in_proj/out_proj (standard GEMMs), and 7 * 3 = 21 for scan (non-GEMM kernel, higher op count per element).
Per-layer FLOPs for Multi-Latent Attention (MLA).
Shared by DeepSeek V3, Kimi K2.5, Mistral Small 4, GLM-5, etc.
When index_topk is set (DSA / sparse attention), accounts for:
- Sparse main attention BMM: S * index_topk instead of 0.5 * S^2
- DSA indexer overhead: Q/K/weights projections + full S^2 indexer BMM
FLOPs for MLA + MoE transformer models (DeepSeek-V3 style).
Parameters:
List of 0/1 per layer (0=dense, 1=MoE).
Total intermediate size for all shared experts combined.
If set, use DSA sparse attention with this many selected positions.
Number of heads in the DSA indexer.
Head dimension of the DSA indexer.
Model FLOPs for MLP layer. Assume gated linear unit.
Model FLOPs for a MoE layer in Nemotron V3/Super V3 (hybrid Mamba/Attention/MoE).
Nemotron V3 uses relu2 (non-gated) for both routed and shared experts, so each expert has 2 linear projections (up_proj + down_proj), not 3.
When moe_latent_size is set (Super V3), routed experts operate in a reduced latent space with additional projection layers (fc1_latent_proj, fc2_latent_proj). The shared expert and gate always operate in the full hidden_size dimension.
Model FLOPs for the Multi-Token-Prediction (MTP) head of Nemotron-3 Super/Ultra.
The head predicts num_mtp_layers (N) additional tokens. Each of the N depths runs
the MTP block pattern once, plus a depth-fusion projection (eh_proj: cat[enorm(embed),
hnorm(hidden)] of size 2*hidden -> hidden) and a vocab projection (weight-tied lm_head).
Repeated layer: when use_repeated_layer is True the model builds a SINGLE physical
depth (mtp_block_types lists its sublayers) and reuses it across all N depths — but
it still EXECUTES once per depth, so the block compute is N x the physical sublayers.
That N x is the repeated layer’s FLOPs. When False, mtp_block_types already spans all
N physical depths, so the block runs once.
These settings are NOT fully recoverable from the HF config (it retains only the physical
depth count and omits the block pattern), so callers pass the effective values read from
the built model: num_mtp_layers = model.mtp_config.num_layers and
mtp_block_types = [s.block_type for s in model.mtp.layers].
Parameters:
effective MTP depths actually run (model.mtp_config.num_layers).
block types (“mamba”/“attention”/“mlp”/“moe”) of the physical MTP sublayers (model.mtp.layers).
True if the physical depth is reused across the N depths.
Model FLOPs for attention layer
sum_{j=1..seq_len} min(cap, floor(j / ratio)) in closed form (cap=None disables the cap).
floor(j / ratio) is the number of complete compressed groups visible to the query at
zero-based position j - 1; the cap is the sparse top-k. Closed form so million-token
sequences cost nothing to evaluate.
sum_{i=0..seq_len-1} min(i + 1, window): causal sliding-window keys per query, summed.
sum_{j=1..seq_len} (j mod ratio) in closed form.
j mod ratio is the length of the incomplete compressed tail visible to the query at
zero-based position j - 1.
Calculate the flops for the attention part.
Model FLOPs for BERT family - accepts either AutoConfig or normalized config
Calculate Model FLOPs Utilization (MFU).
Parameters:
Total TFLOPs across all devices for the measured step.
Total number of GPUs.
Time taken for computation.
Peak TFLOPs/s per device for the training precision. The legacy default is the H100 dense-FP8 peak.
Returns: float
MFU as a percentage.
Model FLOPs for CLIP ViT
Model FLOPs for DeepSeek-V4.1 (CSA2 sparse attention, single-pass mHC, Engram, MoE).
Accepts DeepseekV41TextConfig or the multimodal DeepseekV41Config wrapper (its
text_config is used; the vision tower is not counted, as for other VL entries). The
module shapes follow nemo_automodel/components/models/deepseek_v41:
- attention linears per layer:
wq_a(hidden x q_lora_rank),wq_b(q_lora_rank x headshead_dim),wkv(hidden x head_dim, one shared latent), groupedwo_a(headshead_dim -> o_groupso_lora_rank) andwo_b(o_groupso_lora_rank x hidden); - compressors on
kv_source_layer_ids:wkvand, for ratio > 1,wgate(hidden x head_dim each, applied to every token before pooling); - sparse attention BMMs: every query attends to
min(i+1, sliding_window)local keys plusmin(index_topk, floor((i+1)/ratio))selected compressed keys (SWA-only layers havecompress_ratios[i] == 0); QK^T and PV each costheads * head_dimMACs per key; - the indexer is frozen (
_Indexer.requires_grad_(False); hard top-k has no gradient), so its projections and its causal scoring over the compressed positions are counted forward only (2 x MACs) — the implementation scores the full[S, S/ratio]block before masking, which is executed work but not model work; - MoE per layer: router GEMM (hidden x n_routed_experts) plus
num_experts_per_tok + n_shared_expertsexperts of3 * hidden * moe_intermediate_size; - single-pass mHC: two coefficient projections per layer (
hc_mult*(hc_mult+2)xhc_mult*hidden); the stream collapse/expand are elementwise and not counted; - Engram on
engram_layer_ids: the fused key/value projection ((engram_max_ngram_size-1)*engram_n_heads*engram_head_dimxhidden*(hc_mult+1)); the table gather, hashing, norms, RoPE and quantize/dequantize boundaries are not GEMMs; - the FP32 language-model head (hidden x vocab).
Trainable GEMMs and attention BMMs cost 6 x MACs (forward + backward); activation
recomputation is excluded (that is HFU). For the released Flash configuration the formula
counts 16.09B active GEMM parameters per token (model card: 16B activated during decode)
and 105.6 GFLOPs/token at 4096 tokens, 1.095x the dense-FFN identity 6 x active.
Parameters:
DeepseekV41TextConfig, or a DeepseekV41Config wrapper whose text_config is used.
Number of sequences per step.
Tokens per sequence; defaults to config.max_position_embeddings.
Returns: float
Model FLOPs for one training step of gbs sequences of seq_len tokens.
Model FLOPs for DeepSeek V3 - accepts either AutoConfig or normalized config
Model FLOPs for FLUX
Get the appropriate FLOPs formula function for a given HuggingFace config.
Parameters:
HuggingFace model config object
Returns: Callable | None
The appropriate FLOPs formula function, or None for an unregistered
Estimate FLOPs for GLM4 MoE model configurations.
Model FLOPs for GPT3 family - accepts either AutoConfig or normalized config
Model FLOPs for GPT-OSS
Calculate the flops for the GPT-OSS model
Model FLOPs for Kimi K3 (hybrid KDA / gated-MLA attention + latent MoE with a SiTU shared expert).
Accepts KimiK3TextConfig or the multimodal KimiK3Config wrapper (its
text_config is used; the vision tower is not counted, as for other VL entries).
The layer pattern is read from the config rather than assumed: 1-based
linear_attn_config["kda_layers"] marks KDA layers and every other layer
is MLA (with an extra output-gate projection when mla_use_output_gate);
first_k_dense_replace / moe_layer_freq mark dense vs MoE MLPs, using
the same rule as KimiDecoderLayer. Routed experts run in a latent space
of routed_expert_hidden_size behind shared down/up projections, and the
shared expert is one SiTU MLP of moe_intermediate_size * num_shared_experts.
The router GEMM (hidden_size x num_experts, about 0.6% of the total), norms
and the scalar attention-residual projections are omitted, matching the other
MoE formulas in this module.
The router top-k field on this config is num_experts_per_token (not the
num_experts_per_tok that the generic transformer_flops fallback probes),
which is why K3 was previously costed as a dense 93-layer transformer at well
under half of its real FLOPs. For the released checkpoint’s configuration
(96 attention heads, dense intermediate_size 33792) the formula counts
104.0B active matmul parameters per token and 2.78T total, matching the
model card (“104B activated / 2.8T”) and the tech report’s Table 1 (104.2B /
2.78T) to within 0.2%, which is inside the accounting difference for the
non-GEMM parameters (router, norms, gate biases). The report’s
single MTP layer is not shipped in the checkpoint (num_nextn_predict_layers
is 0) and is not modelled by Automodel, so it is not counted.
Model FLOPs for llama2 family - accepts either AutoConfig or normalized config
Model FLOPs for llama3 family - accepts either AutoConfig or normalized config
Calculate the flops for the loss
Model FLOPs for MiniMax-M2 family - accepts either AutoConfig or normalized config.
Architecture: GQA attention (Q/K/V/O separate projections, head_dim may differ from hidden_size // num_heads) + MoE with SwiGLU (no shared experts by default). Optionally includes MTP (Multi-Token Prediction) modules gated by use_mtp.
Model FLOPs for mixtral family - accepts either AutoConfig or normalized config
Model FLOPs for MLA + MoE models (Kimi K2, GLM-5, Mistral Small 4, etc.).
Handles VL wrappers by extracting text_config if present.
Calculate the flops for the MLP
Model FLOPs for nemotron family - accepts either AutoConfig or normalized config
Model FLOPs for NemotronH
Model FLOPs for NeVA Projection
Model FLOPs for Qwen3.5 family (MoE and Dense) with hybrid GDN/full attention.
Qwen3.5 uses a hybrid attention pattern: 75% GDN (linear attention) layers and 25% standard GQA (full attention) layers (full_attention_interval=4). Supports both the MoE variant (Qwen3.5-35B-A3B) and Dense variant (Qwen3.5-27B).
Model FLOPs for Qwen3.8-Flash-Next (GDN + compressed-block QSA, HyperConnections, Engram PLE, MoE).
Accepts Qwen3_8_FlashNextTextConfig or the multimodal Qwen3_8_FlashNextConfig wrapper
(its text_config is used; the vision tower is not loaded by the training path). The module
shapes follow nemo_automodel/components/models/qwen3_8_flash_next:
- GatedDeltaNet layers (
layers_block_type == "linear_attention"): the shared_gdn_attention_per_layer_flopsterm (QKV/Z/B/A projections, causal conv, chunked delta-rule recurrence, output projection); - QSA full-attention layers: gated
q_proj(hidden x 2headshead_dim),k_proj/v_proj(hidden x kv_heads*head_dim) ando_proj; the sparse GQA BMMs where the query at zero-based positiontattends toratio * min(indexer_budget / ratio, floor((t+1)/ratio))routed tokens plus the(t+1) mod ratiotokens of its incomplete causal tail, QK^T and PV each costingheads * head_dimMACs per key; - the QSA indexer is frozen (
requires_grad_(False); hard top-k has no gradient), so itsindex_qk_projand its causal scoring ofindexer_n_headsqueries against thefloor((t+1)/ratio)compressed keys are counted forward only (2 x MACs); - MoE per layer: router GEMM (hidden x num_experts),
num_experts_per_tokrouted plus one shared expert of3 * hidden * moe_intermediate_size/shared_expert_intermediate_size, and the shared-expert gate (hidden x 1); - HyperConnections: two mixers per layer (attention and MoE), each with
input_mix_weight_down(hc_counthidden x hc_lowrank),input_mix_weight_up(hc_lowrank x hc_counthidden) andblock_inject_weight(hc_count*hidden x hc_count); the final read mixer has no inject weight; - Engram PLE on
ple_layer_ids:key_proj(ple_embed_dim x hc_counthidden),value_proj(ple_embed_dim x hidden) and the depthwise causal convolution (hc_counthidden x kernel); the table gather, hashing and norms are not GEMMs; - the untied language-model head (hidden x vocab). The checkpoint’s MTP head is not loaded.
Trainable GEMMs and attention BMMs cost 6 x MACs (forward + backward); activation recomputation
is excluded (that is HFU). On the released configuration the formula counts 6.14B active GEMM
parameters per token excluding the LM head (model card: 6B activated), implies a 125.8B backbone
excluding the 51.2B Engram table (model card: 125B), and gives 42.0 GFLOPs/token at 4096 tokens,
1.034x the dense identity 6 x active.
Parameters:
Qwen3_8_FlashNextTextConfig, or a Qwen3_8_FlashNextConfig wrapper whose
text_config is used.
Number of sequences per step.
Tokens per sequence; defaults to config.max_position_embeddings.
Returns: float
Model FLOPs for one training step of gbs sequences of seq_len tokens.
Model FLOPs for Qwen3 family - accepts either AutoConfig or normalized config
Model FLOPs for Step3.5-Flash (GQA + sliding-window / full attention + MoE).
Architecture: hybrid full/SWA attention with different head counts per type, MoE with shared expert on most layers, first few layers dense, SwiGLU.
Calculate FLOPs for a standard Transformer model - accepts either AutoConfig or normalized config. Note: This does not cover encoder-decoder models.