> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.llm.benchmark

## Module Contents

### Classes

| Name                                                                                                                         | Description                                    |
| ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- |
| [`BenchmarkingRecipeForNextTokenPrediction`](#nemo_automodel-recipes-llm-benchmark-BenchmarkingRecipeForNextTokenPrediction) | Benchmarking recipe for next-token prediction. |

### Functions

| Name                                                                           | Description                                                                                     |
| ------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- |
| [`_infer_vocab_size`](#nemo_automodel-recipes-llm-benchmark-_infer_vocab_size) | Infer vocab\_size from a model config, handling custom config classes and VL composite configs. |
| [`main`](#nemo_automodel-recipes-llm-benchmark-main)                           | Main entry point for the benchmarking recipe.                                                   |

### Data

[`logger`](#nemo_automodel-recipes-llm-benchmark-logger)

### API

```python
class nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction(
    cfg
)
```

**Bases:** [TrainFinetuneRecipeForNextTokenPrediction](/nemo-automodel/nemo_automodel/recipes/llm/train_ft#nemo_automodel-recipes-llm-train_ft-TrainFinetuneRecipeForNextTokenPrediction)

Benchmarking recipe for next-token prediction.

This class extends TrainFinetuneRecipeForNextTokenPrediction to provide
a simplified benchmarking-focused training loop with timers and profiling support.
It reuses the setup() and \_forward\_backward\_step() methods from the parent class.

benchmark.flops\_scope defaults to "model". Explicit "text" scope counts
only the text backbone over the complete measured iteration time; vision
encoder and projector FLOPs are excluded.

**`_bench_flops_scope`** `= getattr(bench_cfg, 'flops_scope', 'model')`

---

**`_bench_json_output_path`** `= getattr(bench_cfg, 'json_output_path', None)`

---

**`_bench_nsys_end`** `= bench_cfg.nsys_end`

---

**`_bench_nsys_ranks`** `= bench_cfg.nsys_ranks`

---

**`_bench_nsys_start`** `= bench_cfg.nsys_start`

---

**`_bench_peak_tflops`** `= bench_cfg.peak_tflops`

---

**`_bench_seq_len`** `= cfg.dataset.seq_len`

---

**`_bench_steps`** `= cfg.step_scheduler.max_steps`

---

**`_bench_warmup_steps`** `= bench_cfg.warmup_steps`

---

**`_wandb_enabled`** `= cfg.get('wandb', None) is not None`

---

**`timers`** `= Timers(log_level=2, log_option='minmax')`

---

```python
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._log_benchmark_summary(
    steps,
    warmup_steps,
    peak_tflops,
    rank
)
```

```python
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._log_iteration_metrics(
    iter_timer,
    ga_steps,
    peak_tflops,
    rank,
    iteration
)
```

```python
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._mtp_tflops(
    global_batch_size,
    seq_len
)
```

TFLOPs added by a Multi-Token-Prediction (MTP) head, if the model has one.

The backbone FLOPs formula omits the MTP head, and the HF config retains only the
physical depth count (and no per-depth block pattern). So read the EFFECTIVE settings
from the built model: `mtp_config.num_layers` (depths actually run),
`mtp_config.use_repeated_layer`, and the per-sublayer `block_type` from
`mtp.layers`. Returns 0.0 when the model has no enabled MTP head.

```python
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction.run_benchmark()
```

Run the benchmarking loop.

This method implements a simplified training loop focused on benchmarking
with timers and profiling support, similar to the original benchmarking script.

```python
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction.setup()
```

Setup the benchmarking environment.

This method calls the parent's setup() but adapts it for benchmarking purposes.
It skips validation dataloader, checkpointing, and other training-specific features.

```python
nemo_automodel.recipes.llm.benchmark._infer_vocab_size(
    model_cfg
)
```

Infer vocab\_size from a model config, handling custom config classes and VL composite configs.

**Parameters:**

**`model_cfg`**

The model config section (cfg.model) containing *target*, config, etc.

---

**Returns:**

The vocab\_size integer, or raises AttributeError if not found.

```python
nemo_automodel.recipes.llm.benchmark.main(
    config_path = None
)
```

Main entry point for the benchmarking recipe.

Loads the configuration, sets up the recipe, and runs the benchmark.

```python
nemo_automodel.recipes.llm.benchmark.logger = logging.getLogger(__name__)
```