> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.minimax_m3_vl.flops

Useful text-backbone FLOPs for MiniMax M3 training.

## Module Contents

### Functions

| Name                                                                               | Description                                                    |
| ---------------------------------------------------------------------------------- | -------------------------------------------------------------- |
| [`model_flops`](#nemo_automodel-components-models-minimax_m3_vl-flops-model_flops) | Count useful GEMM FLOPs for text-only full-parameter training. |

### API

```python
nemo_automodel.components.models.minimax_m3_vl.flops.model_flops(
    config: nemo_automodel.components.models.minimax_m3_vl.config.MiniMaxM3VLTextConfig,
    gbs: int = 1,
    seq_len: int | None = None
) -> float
```

Count useful GEMM FLOPs for text-only full-parameter training.

Trainable projections and attention count forward, input gradients and weight
gradients (six FLOPs per MAC). Hard block selection has no gradient, so the
indexer projections and causal scores count forward only (two per MAC).
Activation recomputation, masked-out dense work, padding, optimizer updates,
communication, softmax and elementwise operations are excluded. This is MFU,
not the FLOPs actually executed by a particular backend (HFU).

Each sequence is one document without padding. Sparse layers force the
current block into the top-k budget; its causal partial length is counted
exactly. Packed-document FLOPs require the individual document lengths and
cannot be inferred from a packed row length.

**Parameters:**

**`config`** `MiniMaxM3VLTextConfig`

Effective text-backbone configuration, with MTP disabled.

---

**`gbs`** `int` — default: 1

Number of full-length sequences in the global optimizer step.

---

**`seq_len`** `int | None` — default: None

Tokens in each sequence; defaults to max\_position\_embeddings.

---

**Returns:** `float`

Useful model FLOPs over all devices for one optimizer step.

**Raises:**

* `ValueError`: Sequence dimensions, MTP, or sparse selection settings cannot
  be represented by this text-backbone accounting.