> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# nemo_gym.benchmarks

Benchmark discovery and preparation utilities.

## Module Contents

### Classes

| Name                                                      | Description |
| --------------------------------------------------------- | ----------- |
| [`BenchmarkConfig`](#nemo_gym-benchmarks-BenchmarkConfig) | -           |

### Functions

| Name                                                                              | Description                                                                               |
| --------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| [`_benchmark_config_name`](#nemo_gym-benchmarks-_benchmark_config_name)           | The name of the benchmark config, given its path relative to `benchmarks/`, sans `.yaml`. |
| [`_benchmark_config_paths`](#nemo_gym-benchmarks-_benchmark_config_paths)         | Sorted config paths under one dir that declare a benchmark, discovered by content.        |
| [`_discover_benchmarks_in_dir`](#nemo_gym-benchmarks-_discover_benchmarks_in_dir) | Map benchmark name -> `BenchmarkConfig` for every benchmark config under one dir.         |
| [`_is_benchmark_config`](#nemo_gym-benchmarks-_is_benchmark_config)               | True if the config declares exactly one `type: benchmark` dataset in its own structure.   |
| [`discover_benchmarks`](#nemo_gym-benchmarks-discover_benchmarks)                 | Map benchmark name -> `BenchmarkConfig` for every discoverable benchmark config.          |

### Data

[`BENCHMARKS_DIR`](#nemo_gym-benchmarks-BENCHMARKS_DIR)

[`BENCHMARKS_SUBDIR`](#nemo_gym-benchmarks-BENCHMARKS_SUBDIR)

[`MANIFEST_FILENAME`](#nemo_gym-benchmarks-MANIFEST_FILENAME)

[`__getattr__`](#nemo_gym-benchmarks-__getattr__)

### API

```python
class nemo_gym.benchmarks.BenchmarkConfig()
```

**Bases:** `BaseModel`

**`agent_name`** `str`

---

**`dataset`** `BenchmarkDatasetConfig`

---

**`name`** `str`

---

**`num_repeats`** `int`

---

**`path`** `Path`

---

```python
nemo_gym.benchmarks.BenchmarkConfig.from_config_path(
    config_path: pathlib.Path,
    strict: bool = True
) -> typing.Optional[nemo_gym.benchmarks.BenchmarkConfig]
```

classmethod

```python
nemo_gym.benchmarks.BenchmarkConfig.from_initial_config_dict(
    path: pathlib.Path,
    initial_config_dict: omegaconf.DictConfig,
    strict: bool = True
) -> typing.Optional[nemo_gym.benchmarks.BenchmarkConfig]
```

classmethod

```python
nemo_gym.benchmarks._benchmark_config_name(
    rel_config_path: pathlib.Path
) -> str
```

The name of the benchmark config, given its path relative to `benchmarks/`, sans `.yaml`.

This is the identity we key benchmarks by, so a listed benchmark is always a valid `--benchmark` argument.

```python
nemo_gym.benchmarks._benchmark_config_paths(
    benchmarks_dir: pathlib.Path
) -> typing.List[pathlib.Path]
```

Sorted config paths under one dir that declare a benchmark, discovered by content.

A config is a benchmark iff it declares exactly one `type: benchmark` dataset, regardless of filename, so
we scan every yaml. `_is_benchmark_config` is a cheap prefilter (pay the resolve cost only on real
candidates) that also catches non-`config.yaml` names like tau2's `configs/*.yaml`. Empty if dir missing.

```python
nemo_gym.benchmarks._discover_benchmarks_in_dir(
    benchmarks_dir: pathlib.Path
) -> typing.Dict[str, nemo_gym.benchmarks.BenchmarkConfig]
```

Map benchmark name -> `BenchmarkConfig` for every benchmark config under one dir.

```python
nemo_gym.benchmarks._is_benchmark_config(
    config_path: pathlib.Path
) -> bool
```

True if the config declares exactly one `type: benchmark` dataset in its own structure.

A raw single-file parse (no `config_paths`/interpolation resolution), so it's format-agnostic and can't
fail on includes. Declaring several makes it an eval suite rather than a benchmark: there is no single
name, agent, or repeat count to catalog it under, so it is not a valid `--benchmark` argument. An
unparseable file is kept (returns True) so the resolve step surfaces a diagnostic.

```python
nemo_gym.benchmarks.discover_benchmarks() -> typing.Dict[str, nemo_gym.benchmarks.BenchmarkConfig]
```

Map benchmark name -> `BenchmarkConfig` for every discoverable benchmark config.

Scans the `benchmarks/` subdir of every `component_search_roots` root
(`NEMO_GYM_EXTRA_ROOTS` + cwd + built-ins), merged so user benchmarks shadow same-named built-ins.

```python
nemo_gym.benchmarks.BENCHMARKS_DIR = PARENT_DIR / BENCHMARKS_SUBDIR
```

```python
nemo_gym.benchmarks.BENCHMARKS_SUBDIR = 'benchmarks'
```

```python
nemo_gym.benchmarks.MANIFEST_FILENAME = 'manifest.yaml'
```

```python
nemo_gym.benchmarks.__getattr__ = moved_attr_getter(__name__, {'list_benchmarks': 'nemo_gym.cli.eval', 'PrepareBen...
```