> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# nemo_gym.reward_profile

## Module Contents

### Classes

| Name                                                                      | Description                                                             |
| ------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| [`AggregateMetricsMixin`](#nemo_gym-reward_profile-AggregateMetricsMixin) | Mixin providing full-run, per-repeat, and key-metric aggregation hooks. |
| [`RewardProfileConfig`](#nemo_gym-reward_profile-RewardProfileConfig)     | -                                                                       |
| [`RewardProfiler`](#nemo_gym-reward_profile-RewardProfiler)               | -                                                                       |

### Functions

| Name                                                                                      | Description                                                                                  |
| ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| [`_add_custom_repeat_metrics`](#nemo_gym-reward_profile-_add_custom_repeat_metrics)       | Recompute benchmark metrics per repeat, replacing generic collisions and their aggregates.   |
| [`_coverage_metrics`](#nemo_gym-reward_profile-_coverage_metrics)                         | How much of the run the quality metrics above were actually computed from.                   |
| [`_group_by_task`](#nemo_gym-reward_profile-_group_by_task)                               | Group verify responses by task index, returning a list of per-task rollout lists.            |
| [`_is_finite_number`](#nemo_gym-reward_profile-_is_finite_number)                         | -                                                                                            |
| [`_partition_on_mask`](#nemo_gym-reward_profile-_partition_on_mask)                       | Split into the samples whose reward is a valid measurement, and the rest.                    |
| [`_rollout_key`](#nemo_gym-reward_profile-_rollout_key)                                   | -                                                                                            |
| [`_stat`](#nemo_gym-reward_profile-_stat)                                                 | -                                                                                            |
| [`add_avg_sample_std_dev`](#nemo_gym-reward_profile-add_avg_sample_std_dev)               | Add avg\_sample\_std\_dev statistics to an existing metrics dict.                            |
| [`compute_aggregate_metrics`](#nemo_gym-reward_profile-compute_aggregate_metrics)         | Shared aggregation logic for /aggregate\_metrics.                                            |
| [`compute_pass_majority_metrics`](#nemo_gym-reward_profile-compute_pass_majority_metrics) | Compute pass\@k, majority\@k, no\_answer, and variance statistics from grouped task results. |
| [`compute_perf_summary`](#nemo_gym-reward_profile-compute_perf_summary)                   | Aggregate per-rollout `ng_perf` dicts into the `perf_summary` block (RFC R3).                |
| [`compute_subset_metrics`](#nemo_gym-reward_profile-compute_subset_metrics)               | Group tasks by a field and compute pass\@k metrics per subset.                               |
| [`coverage_by_agent`](#nemo_gym-reward_profile-coverage_by_agent)                         | Coverage for each agent, computed from that agent's own records.                             |
| [`highest_k_metrics`](#nemo_gym-reward_profile-highest_k_metrics)                         | Select the highest-k entries matching a metric pattern.                                      |
| [`select_measured`](#nemo_gym-reward_profile-select_measured)                             | Narrow a rollout set to what quality metrics may be computed from.                           |

### Data

[`MASK_SAMPLE_FIELD`](#nemo_gym-reward_profile-MASK_SAMPLE_FIELD)

[`_PERF_SUMMARY_LATENCY_STATS`](#nemo_gym-reward_profile-_PERF_SUMMARY_LATENCY_STATS)

[`_PERF_SUMMARY_NUM_TURNS_STATS`](#nemo_gym-reward_profile-_PERF_SUMMARY_NUM_TURNS_STATS)

[`_PERF_SUMMARY_TOKEN_FIELDS`](#nemo_gym-reward_profile-_PERF_SUMMARY_TOKEN_FIELDS)

[`_PERF_SUMMARY_TOKEN_STATS`](#nemo_gym-reward_profile-_PERF_SUMMARY_TOKEN_STATS)

[`__getattr__`](#nemo_gym-reward_profile-__getattr__)

### API

```python
class nemo_gym.reward_profile.AggregateMetricsMixin()
```

Mixin providing full-run, per-repeat, and key-metric aggregation hooks.

Inherited by both SimpleResourcesServer and SimpleResponsesAPIAgent so that
benchmark-specific metric logic can live on either server type.

```python
nemo_gym.reward_profile.AggregateMetricsMixin.compute_metrics(
    tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]]
) -> typing.Dict[str, typing.Any]
```

Override to compute custom metrics from all verify responses.

Receives verify responses grouped by task: tasks\[i] is a list of rollout
dicts for task i. Each dict has at minimum reward, plus any custom fields
from the verify response (e.g. symbolic\_correct, judgement-gen-base). This
hook is called once for the full dataset.

Use for metrics that need the full dataset at once:

* Confidence intervals (ArenaMetrics)
* Cross-task statistics (std\_dev\_across\_runs)
* pass\@k with proper combinatorial computation

The returned dict is merged into agent\_metrics.
Default: empty dict (no additional metrics).

```python
nemo_gym.reward_profile.AggregateMetricsMixin.compute_repeat_metrics(
    tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]]
) -> typing.Dict[str, typing.Any]
```

Override to compute custom metrics independently for one repeat.

Returned names replace matching generic repeat metrics. A metric is summarized
only when it is finite for every scored repeat and its full-run value, if any, is finite.

```python
nemo_gym.reward_profile.AggregateMetricsMixin.get_key_metrics(
    agent_metrics: typing.Dict[str, typing.Any]
) -> typing.Dict[str, typing.Any]
```

Override to select headline metrics for this benchmark.

Default: all mean/\* entries from agent\_metrics.

```python
class nemo_gym.reward_profile.RewardProfileConfig()
```

**Bases:** [BaseNeMoGymCLIConfig](/nemo/gym/nemo-gym/nemo_gym/config_types#nemo_gym-config_types-BaseNeMoGymCLIConfig)

**`allow_partial_rollouts`** `bool`

---

**`materialized_inputs_jsonl_fpath`** `str`

---

**`rollouts_jsonl_fpath`** `str`

---

```python
class nemo_gym.reward_profile.RewardProfiler()
```

```python
nemo_gym.reward_profile.RewardProfiler._aggregate_repeat_level_metrics(
    repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]]
) -> typing.List[typing.Dict[str, typing.Any]]
```

Aggregate per-repeat estimates (e.g. mean/reward) across repeats, per agent.

Treats each repeat's stat as one observation and reports the statistics across
repeats. Existing summaries and uncertainty estimates are retained in the input
rows but excluded here to avoid calculating statistics over statistics.

```python
nemo_gym.reward_profile.RewardProfiler._append_repeat_level_aggregates_to_agent_level_metrics(
    agent_level_metrics: typing.List[typing.Dict[str, typing.Any]],
    repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]]
) -> None
```

```python
nemo_gym.reward_profile.RewardProfiler._compute_repeat_level_metrics(
    df: pandas.DataFrame
) -> typing.List[typing.Dict[str, typing.Any]]
```

Per-agent, per-rollout-index summary stats across all tasks.

Only produced for agents that have more than one rollout index — agents with a
single rollout contribute nothing to a repeat-level comparison and are skipped.
If no agent qualifies, returns an empty list.

```python
nemo_gym.reward_profile.RewardProfiler._confidence_interval(
    mean: float,
    sem: float,
    n: int,
    confidence: float = 0.95
) -> typing.Optional[typing.Tuple[float, float]]
```

Return (ci\_low, ci\_high) t-interval at the given confidence level, or None when n \<= 1.

```python
nemo_gym.reward_profile.RewardProfiler._index_by_rollout_key(
    rows: typing.List[typing.Dict[str, typing.Any]],
    name: str
) -> typing.Dict[typing.Tuple[int, int], typing.Dict[str, typing.Any]]
```

```python
nemo_gym.reward_profile.RewardProfiler.align_rows_and_results(
    rows: typing.List[typing.Dict[str, typing.Any]],
    results: typing.List[typing.Dict[str, typing.Any]],
    allow_partial_rollouts: bool = False
) -> typing.List[typing.Tuple[typing.Dict[str, typing.Any], typing.Dict[str, typing.Any]]]
```

```python
nemo_gym.reward_profile.RewardProfiler.calculate_metrics_single_df(
    grouped_df: pandas.core.groupby.generic.DataFrameGroupBy
) -> typing.List[typing.Dict[str, typing.Any]]
```

```python
nemo_gym.reward_profile.RewardProfiler.describe_dataframe(
    df: pandas.DataFrame
) -> pandas.DataFrame
```

```python
nemo_gym.reward_profile.RewardProfiler.histogram(
    data: pandas.Series
) -> typing.Optional['Histogram']
```

```python
nemo_gym.reward_profile.RewardProfiler.prepare_for_serialization(
    metrics: typing.List[typing.Dict]
) -> typing.List[typing.Dict]
```

Non-destructively cleans metrics output by RewardProfiler for downstream serialization.

```python
nemo_gym.reward_profile.RewardProfiler.profile_completion_summary(
    rows: typing.List[typing.Dict[str, typing.Any]],
    results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Any]
```

```python
nemo_gym.reward_profile.RewardProfiler.profile_from_data(
    rows: typing.List[typing.Dict[str, typing.Any]],
    results: typing.List[typing.Dict[str, typing.Any]],
    allow_partial_rollouts: bool = False
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]]]
```

```python
nemo_gym.reward_profile.RewardProfiler.rollout_info_from_result(
    result: typing.Dict[str, typing.Any]
) -> typing.Dict[str, typing.Any]
```

```python
nemo_gym.reward_profile.RewardProfiler.write_to_disk(
    group_level_metrics: typing.List[typing.Dict[str, typing.Any]],
    agent_level_metrics: typing.List[typing.Dict[str, typing.Any]],
    repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]],
    base_output_fpath: pathlib.Path
) -> typing.Tuple[pathlib.Path, pathlib.Path, pathlib.Path]
```

```python
nemo_gym.reward_profile._add_custom_repeat_metrics(
    profiler: nemo_gym.reward_profile.RewardProfiler,
    verify_responses: typing.List[typing.Dict[str, typing.Any]],
    repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]],
    agent_metrics: typing.Dict[str, typing.Any],
    custom_metrics: typing.Dict[str, typing.Any],
    compute_repeat_metrics_fn: typing.Any
) -> None
```

Recompute benchmark metrics per repeat, replacing generic collisions and their aggregates.

```python
nemo_gym.reward_profile._coverage_metrics(
    all_responses: typing.List[typing.Dict[str, typing.Any]],
    scored: typing.List[typing.Dict[str, typing.Any]],
    masked: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Any]
```

How much of the run the quality metrics above were actually computed from.

Empty when nothing was masked, so a run that reports no masking publishes exactly the
keys it published before.

```python
nemo_gym.reward_profile._group_by_task(
    verify_responses: typing.List[typing.Dict[str, typing.Any]]
) -> typing.List[typing.List[typing.Dict[str, typing.Any]]]
```

Group verify responses by task index, returning a list of per-task rollout lists.

```python
nemo_gym.reward_profile._is_finite_number(
    value: typing.Any
) -> bool
```

```python
nemo_gym.reward_profile._partition_on_mask(
    verify_responses: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]]]
```

Split into the samples whose reward is a valid measurement, and the rest.

```python
nemo_gym.reward_profile._rollout_key(
    row: typing.Dict[str, typing.Any]
) -> typing.Tuple[typing.Any, typing.Any]
```

```python
nemo_gym.reward_profile._stat(
    values: typing.List[float],
    stat: str
) -> float
```

```python
nemo_gym.reward_profile.add_avg_sample_std_dev(
    metrics: typing.Dict[str, typing.Any],
    all_score_dicts: typing.List[typing.List[typing.Dict[str, float]]],
    score_names: list,
    max_k: int
) -> None
```

Add avg\_sample\_std\_dev statistics to an existing metrics dict.

Computes the average of per-task standard deviations across k rollouts — a measure of
within-task variance that complements the across-run variance (std\_dev\_across\_runs).

Modifies `metrics` in place.

```python
nemo_gym.reward_profile.compute_aggregate_metrics(
    verify_responses: typing.List[typing.Dict[str, typing.Any]],
    compute_metrics_fn = None,
    get_key_metrics_fn = None,
    compute_repeat_metrics_fn = None
) -> nemo_gym.config_types.AggregateMetrics
```

Shared aggregation logic for /aggregate\_metrics.

RewardProfiler runs with defaults to produce baseline stats (mean/max/min/median/std)
for both group-level (per-task) and agent-level metrics.

Optionally accepts custom functions for benchmark-specific customization:

* compute\_metrics\_fn: receives all verify responses grouped by task. Its returned
  dict is merged into agent\_metrics.
* get\_key\_metrics\_fn: select headline metrics from agent\_metrics
* compute\_repeat\_metrics\_fn: receives one repeat grouped by task. Its returned
  metrics are summarized across repeats.

```python
nemo_gym.reward_profile.compute_pass_majority_metrics(
    tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]],
    score_fn: typing.Optional[typing.Any] = None,
    answer_key: typing.Optional[str] = None
) -> typing.Tuple[typing.Dict[str, typing.Any], typing.List[typing.List[typing.Dict[str, float]]], typing.List[str], int]
```

Compute pass\@k, majority\@k, no\_answer, and variance statistics from grouped task results.

Shared utility for any resource server's compute\_metrics() override.

**Parameters:**

**`tasks`** `List[List[Dict[str, Any]]]`

tasks\[i] is a list of rollout dicts for task i.

---

**`score_fn`** `Optional[Any]` — default: None

Callable(result\_dict) -> Dict\[str, float|bool] returning named scores.
Defaults to `lambda r: &#123;"accuracy": r["reward"]&#125;`.

---

**`answer_key`** `Optional[str]` — default: None

Field name for extracted answer (enables majority\@k and no\_answer).
If None, majority\@k and no\_answer are skipped.

---

**Returns:** `Dict[str, Any]`

Metrics, all\_score\_dicts, score\_names, max\_k

```python
nemo_gym.reward_profile.compute_perf_summary(
    ng_perf_records: typing.List[typing.Dict[str, typing.Any]],
    total_rollouts: int
) -> typing.Optional[typing.Dict[str, typing.Any]]
```

Aggregate per-rollout `ng_perf` dicts into the `perf_summary` block (RFC R3).

`total_rollouts` is the full rollout count for this batch, including rollouts that never
produced `ng_perf` at all -- the denominator for `overall_observability_coverage` and
`token_observability_coverage`.

Returns `None` when there are no rollouts at all, or when none of them carried `ng_perf`
(observability off for the whole run, or on but nothing was collected) -- perf\_summary is
absent rather than present with a 0.0 coverage either way; a coverage value, when present, is
always > 0. Every other stat is included only when at least one rollout reported the
underlying field (a provider that never reports cache usage yields no
`mean_cached_prompt_tokens`, for example).

```python
nemo_gym.reward_profile.compute_subset_metrics(
    tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]],
    subset_key: str,
    score_fn: typing.Optional[typing.Any] = None,
    answer_key: typing.Optional[str] = None
) -> typing.Dict[str, typing.Any]
```

Group tasks by a field and compute pass\@k metrics per subset.

Returns flat dict with subset-prefixed keys, e.g. `"easy/pass@1/accuracy"`.
Skips the `per_sample_aggregate` key from each subset's metrics.

**Parameters:**

**`tasks`** `List[List[Dict[str, Any]]]`

tasks\[i] is a list of rollout dicts for task i.

---

**`subset_key`** `str`

Field name in rollout dicts to group by (e.g. `"difficulty"`).

---

**`score_fn`** `Optional[Any]` — default: None

Passed through to `compute_pass_majority_metrics`.

---

**`answer_key`** `Optional[str]` — default: None

Passed through to `compute_pass_majority_metrics`.

---

```python
nemo_gym.reward_profile.coverage_by_agent(
    rows: typing.List[typing.Dict[str, typing.Any]],
    results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Dict[str, typing.Any]]
```

Coverage for each agent, computed from that agent's own records.

A run-wide coverage block copied onto every agent tells each of them how much *the run*
masked, which reads as that agent's own loss. An agent that masked nothing would carry
another agent's count.

Agents are keyed by name; an agent whose every result was masked still gets an entry,
so a caller can keep reporting it after the quality metrics drop it.

```python
nemo_gym.reward_profile.highest_k_metrics(
    agent_metrics: typing.Dict[str, typing.Any],
    pattern: str,
    score_names: typing.Optional[typing.List[str]] = None,
    exclude_names: typing.Optional[typing.List[str]] = None
) -> typing.Dict[str, typing.Any]
```

Select the highest-k entries matching a metric pattern.

Finds all keys matching `pattern` (with `&#123;k&#125;` as the k placeholder), determines the
highest k value, and returns all entries at that k.

Example::

# Get highest-k pass\@k for accuracy only

highest\_k\_metrics(am, "pass@\{k}", score\_names=\["accuracy"])

# → \{"pass\@32/accuracy": 95.0}

# Get highest-k pass\@1\[avg-of-k] for all scores except no\_answer, without stats

highest\_k\_metrics(am, "pass\@1\[avg-of-\{k}]", exclude\_names=\["no\_answer"])

# → \{"pass\@1\[avg-of-32]/accuracy": 94.5, "pass\@1\[avg-of-32]/symbolic\_accuracy": 93.2}

**Parameters:**

**`agent_metrics`** `Dict[str, Any]`

Full agent metrics dict.

---

**`pattern`** `str`

Pattern with `&#123;k&#125;` placeholder, e.g. `"pass@&#123;k&#125;"` or `"pass@1[avg-of-&#123;k&#125;]"`.

---

**`score_names`** `Optional[List[str]]` — default: None

If provided, only return entries whose score name (after the last `/`)
is in this list. Stat suffixes (std\_dev, std\_err, avg\_sample) are always excluded.

---

**`exclude_names`** `Optional[List[str]]` — default: None

Score names to exclude (e.g. `["no_answer"]`). Applied after score\_names.

---

**Returns:** `Dict[str, Any]`

Dict of matching metrics at the highest k, e.g. `&#123;"pass@32/accuracy": 95.0&#125;`.

```python
nemo_gym.reward_profile.select_measured(
    rows: typing.List[typing.Dict[str, typing.Any]],
    results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.Dict[str, typing.Any]]
```

Narrow a rollout set to what quality metrics may be computed from.

Both the aggregation path and `gym eval profile` go through here, so the same saved
rollouts produce the same quality numbers whichever view you look at.

Three things happen. Masked samples leave the quality set, because their reward is not
a measurement of the evaluated system. The flag itself is stripped from what remains:
it is a flag, not a measurement, and the profiler would otherwise coerce it to an int
and publish `mean/mask_sample` alongside real metrics. And the input rows of exactly
those masked results are dropped with them, so the two stay aligned -- a masked rollout
is absent from the quality set but was never missing from the collection, and must not
be reported as an incomplete run or require `allow_partial_rollouts` to profile.

Only the masked pairs are removed, never "keep what was scored": a row whose result is
genuinely missing has to survive into the quality set so alignment still reports the
collection as partial. This function narrows what is measured; it does not decide
whether the collection was complete, and callers that enforce completeness must
validate the original rows and results before calling it.

Masked rows are returned rather than discarded: completion and coverage accounting
still has to see them.

```python
nemo_gym.reward_profile.MASK_SAMPLE_FIELD = 'mask_sample'
```

```python
nemo_gym.reward_profile._PERF_SUMMARY_LATENCY_STATS: Tuple[str, ...] = ('p50', 'p90', 'p99', Stat.MEAN)
```

```python
nemo_gym.reward_profile._PERF_SUMMARY_NUM_TURNS_STATS: Tuple[str, ...] = (Stat.MEAN, Stat.MEDIAN, Stat.STD, Stat.MIN, 'p90', Stat.MAX)
```

```python
nemo_gym.reward_profile._PERF_SUMMARY_TOKEN_FIELDS: Tuple[str, ...] = ('prompt_tokens', 'cached_prompt_tokens', 'completion_tokens', 'reasoning_tokens...
```

```python
nemo_gym.reward_profile._PERF_SUMMARY_TOKEN_STATS: Tuple[str, ...] = (Stat.MEAN, Stat.MEDIAN, 'total')
```

```python
nemo_gym.reward_profile.__getattr__ = moved_attr_getter(__name__, {'reward_profile': 'nemo_gym.cli.eval'})
```