nemo_automodel.recipes.llm.benchmark

View as Markdown

Module Contents

Classes

NameDescription
BenchmarkingRecipeForNextTokenPredictionBenchmarking recipe for next-token prediction.

Functions

NameDescription
_infer_vocab_sizeInfer vocab_size from a model config, handling custom config classes and VL composite configs.
mainMain entry point for the benchmarking recipe.

Data

logger

API

class nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction(
cfg
)

Bases: TrainFinetuneRecipeForNextTokenPrediction

Benchmarking recipe for next-token prediction.

This class extends TrainFinetuneRecipeForNextTokenPrediction to provide a simplified benchmarking-focused training loop with timers and profiling support. It reuses the setup() and _forward_backward_step() methods from the parent class.

benchmark.flops_scope defaults to “model”. Explicit “text” scope counts only the text backbone over the complete measured iteration time; vision encoder and projector FLOPs are excluded.

_bench_flops_scope
= getattr(bench_cfg, 'flops_scope', 'model')
_bench_json_output_path
= getattr(bench_cfg, 'json_output_path', None)
_bench_nsys_end
= bench_cfg.nsys_end
_bench_nsys_ranks
= bench_cfg.nsys_ranks
_bench_nsys_start
= bench_cfg.nsys_start
_bench_peak_tflops
= bench_cfg.peak_tflops
_bench_seq_len
= cfg.dataset.seq_len
_bench_steps
= cfg.step_scheduler.max_steps
_bench_warmup_steps
= bench_cfg.warmup_steps
_wandb_enabled
= cfg.get('wandb', None) is not None
timers
= Timers(log_level=2, log_option='minmax')
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._log_benchmark_summary(
steps,
warmup_steps,
peak_tflops,
rank
)
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._log_iteration_metrics(
iter_timer,
ga_steps,
peak_tflops,
rank,
iteration
)
nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction._mtp_tflops(
global_batch_size,
seq_len
)

TFLOPs added by a Multi-Token-Prediction (MTP) head, if the model has one.

The backbone FLOPs formula omits the MTP head, and the HF config retains only the physical depth count (and no per-depth block pattern). So read the EFFECTIVE settings from the built model: mtp_config.num_layers (depths actually run), mtp_config.use_repeated_layer, and the per-sublayer block_type from mtp.layers. Returns 0.0 when the model has no enabled MTP head.

nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction.run_benchmark()

Run the benchmarking loop.

This method implements a simplified training loop focused on benchmarking with timers and profiling support, similar to the original benchmarking script.

nemo_automodel.recipes.llm.benchmark.BenchmarkingRecipeForNextTokenPrediction.setup()

Setup the benchmarking environment.

This method calls the parent’s setup() but adapts it for benchmarking purposes. It skips validation dataloader, checkpointing, and other training-specific features.

nemo_automodel.recipes.llm.benchmark._infer_vocab_size(
model_cfg
)

Infer vocab_size from a model config, handling custom config classes and VL composite configs.

Parameters:

model_cfg

The model config section (cfg.model) containing target, config, etc.

Returns:

The vocab_size integer, or raises AttributeError if not found.

nemo_automodel.recipes.llm.benchmark.main(
config_path = None
)

Main entry point for the benchmarking recipe.

Loads the configuration, sets up the recipe, and runs the benchmark.

nemo_automodel.recipes.llm.benchmark.logger = logging.getLogger(__name__)