nemo_gym.reward_profile

View as Markdown

Module Contents

Classes

NameDescription
AggregateMetricsMixinMixin providing compute_metrics/get_key_metrics hooks and the aggregate_metrics endpoint.
RewardProfileConfig-
RewardProfiler-

Functions

NameDescription
_coverage_metricsHow much of the run the quality metrics above were actually computed from.
_group_by_taskGroup verify responses by task index, returning a list of per-task rollout lists.
_partition_on_maskSplit into the samples whose reward is a valid measurement, and the rest.
_rollout_key-
_stat-
add_avg_sample_std_devAdd avg_sample_std_dev statistics to an existing metrics dict.
compute_aggregate_metricsShared aggregation logic for /aggregate_metrics.
compute_pass_majority_metricsCompute pass@k, majority@k, no_answer, and variance statistics from grouped task results.
compute_perf_summaryAggregate per-rollout ng_perf dicts into the perf_summary block (RFC R3).
compute_subset_metricsGroup tasks by a field and compute pass@k metrics per subset.
coverage_by_agentCoverage for each agent, computed from that agent’s own records.
highest_k_metricsSelect the highest-k entries matching a metric pattern.
select_measuredNarrow a rollout set to what quality metrics may be computed from.

Data

MASK_SAMPLE_FIELD

_PERF_SUMMARY_LATENCY_STATS

_PERF_SUMMARY_NUM_TURNS_STATS

_PERF_SUMMARY_TOKEN_FIELDS

_PERF_SUMMARY_TOKEN_STATS

__getattr__

API

class nemo_gym.reward_profile.AggregateMetricsMixin()

Mixin providing compute_metrics/get_key_metrics hooks and the aggregate_metrics endpoint.

Inherited by both SimpleResourcesServer and SimpleResponsesAPIAgent so that benchmark-specific metric logic can live on either server type.

nemo_gym.reward_profile.AggregateMetricsMixin.compute_metrics(
tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]]
) -> typing.Dict[str, typing.Any]

Override to compute custom metrics from all verify responses.

Receives verify responses grouped by task: tasks[i] is a list of rollout dicts for task i. Each dict has at minimum reward, plus any custom fields from the verify response (e.g. symbolic_correct, judgement-gen-base).

Use for metrics that need the full dataset at once:

  • Confidence intervals (ArenaMetrics)
  • Cross-task statistics (std_dev_across_runs)
  • pass@k with proper combinatorial computation

The returned dict is merged into agent_metrics. Default: empty dict (no additional metrics).

nemo_gym.reward_profile.AggregateMetricsMixin.get_key_metrics(
agent_metrics: typing.Dict[str, typing.Any]
) -> typing.Dict[str, typing.Any]

Override to select headline metrics for this benchmark.

Default: all mean/* entries from agent_metrics.

class nemo_gym.reward_profile.RewardProfileConfig()

Bases: BaseNeMoGymCLIConfig

allow_partial_rollouts
bool
materialized_inputs_jsonl_fpath
str
rollouts_jsonl_fpath
str
class nemo_gym.reward_profile.RewardProfiler()
nemo_gym.reward_profile.RewardProfiler._aggregate_repeat_level_metrics(
repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]]
) -> typing.List[typing.Dict[str, typing.Any]]

Aggregate per-repeat estimates (e.g. mean/reward) across repeats, per agent.

Treats each repeat’s stat as one observation and reports the statistics across repeats.

nemo_gym.reward_profile.RewardProfiler._append_repeat_level_aggregates_to_agent_level_metrics(
agent_level_metrics: typing.List[typing.Dict[str, typing.Any]],
repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]]
) -> None
nemo_gym.reward_profile.RewardProfiler._compute_repeat_level_metrics(
df: pandas.DataFrame
) -> typing.List[typing.Dict[str, typing.Any]]

Per-agent, per-rollout-index summary stats across all tasks.

Only produced for agents that have more than one rollout index — agents with a single rollout contribute nothing to a repeat-level comparison and are skipped. If no agent qualifies, returns an empty list.

nemo_gym.reward_profile.RewardProfiler._confidence_interval(
mean: float,
sem: float,
n: int,
confidence: float = 0.95
) -> typing.Optional[typing.Tuple[float, float]]

Return (ci_low, ci_high) t-interval at the given confidence level, or None when n <= 1.

nemo_gym.reward_profile.RewardProfiler._index_by_rollout_key(
rows: typing.List[typing.Dict[str, typing.Any]],
name: str
) -> typing.Dict[typing.Tuple[int, int], typing.Dict[str, typing.Any]]
nemo_gym.reward_profile.RewardProfiler.align_rows_and_results(
rows: typing.List[typing.Dict[str, typing.Any]],
results: typing.List[typing.Dict[str, typing.Any]],
allow_partial_rollouts: bool = False
) -> typing.List[typing.Tuple[typing.Dict[str, typing.Any], typing.Dict[str, typing.Any]]]
nemo_gym.reward_profile.RewardProfiler.calculate_metrics_single_df(
grouped_df: pandas.core.groupby.generic.DataFrameGroupBy
) -> typing.List[typing.Dict[str, typing.Any]]
nemo_gym.reward_profile.RewardProfiler.describe_dataframe(
df: pandas.DataFrame
) -> pandas.DataFrame
nemo_gym.reward_profile.RewardProfiler.histogram(
data: pandas.Series
) -> typing.Optional[wandb.Histogram]
nemo_gym.reward_profile.RewardProfiler.prepare_for_serialization(
metrics: typing.List[typing.Dict]
) -> typing.List[typing.Dict]

Non-destructively cleans metrics output by RewardProfiler for downstream serialization.

nemo_gym.reward_profile.RewardProfiler.profile_completion_summary(
rows: typing.List[typing.Dict[str, typing.Any]],
results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Any]
nemo_gym.reward_profile.RewardProfiler.profile_from_data(
rows: typing.List[typing.Dict[str, typing.Any]],
results: typing.List[typing.Dict[str, typing.Any]],
allow_partial_rollouts: bool = False
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]]]
nemo_gym.reward_profile.RewardProfiler.rollout_info_from_result(
result: typing.Dict[str, typing.Any]
) -> typing.Dict[str, typing.Any]
nemo_gym.reward_profile.RewardProfiler.write_to_disk(
group_level_metrics: typing.List[typing.Dict[str, typing.Any]],
agent_level_metrics: typing.List[typing.Dict[str, typing.Any]],
repeat_level_metrics: typing.List[typing.Dict[str, typing.Any]],
base_output_fpath: pathlib.Path
) -> typing.Tuple[pathlib.Path, pathlib.Path, pathlib.Path]
nemo_gym.reward_profile._coverage_metrics(
all_responses: typing.List[typing.Dict[str, typing.Any]],
scored: typing.List[typing.Dict[str, typing.Any]],
masked: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Any]

How much of the run the quality metrics above were actually computed from.

Empty when nothing was masked, so a run that reports no masking publishes exactly the keys it published before.

nemo_gym.reward_profile._group_by_task(
verify_responses: typing.List[typing.Dict[str, typing.Any]]
) -> typing.List[typing.List[typing.Dict[str, typing.Any]]]

Group verify responses by task index, returning a list of per-task rollout lists.

nemo_gym.reward_profile._partition_on_mask(
verify_responses: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]]]

Split into the samples whose reward is a valid measurement, and the rest.

nemo_gym.reward_profile._rollout_key(
row: typing.Dict[str, typing.Any]
) -> typing.Tuple[typing.Any, typing.Any]
nemo_gym.reward_profile._stat(
values: typing.List[float],
stat: str
) -> float
nemo_gym.reward_profile.add_avg_sample_std_dev(
metrics: typing.Dict[str, typing.Any],
all_score_dicts: typing.List[typing.List[typing.Dict[str, float]]],
score_names: list,
max_k: int
) -> None

Add avg_sample_std_dev statistics to an existing metrics dict.

Computes the average of per-task standard deviations across k rollouts — a measure of within-task variance that complements the across-run variance (std_dev_across_runs).

Modifies metrics in place.

nemo_gym.reward_profile.compute_aggregate_metrics(
verify_responses: typing.List[typing.Dict[str, typing.Any]],
compute_metrics_fn = None,
get_key_metrics_fn = None

Shared aggregation logic for /aggregate_metrics.

RewardProfiler runs with defaults to produce baseline stats (mean/max/min/median/std) for both group-level (per-task) and agent-level metrics.

nemo_gym.reward_profile.compute_pass_majority_metrics(
tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]],
score_fn: typing.Optional[typing.Any] = None,
answer_key: typing.Optional[str] = None
) -> typing.Tuple[typing.Dict[str, typing.Any], typing.List[typing.List[typing.Dict[str, float]]], typing.List[str], int]

Compute pass@k, majority@k, no_answer, and variance statistics from grouped task results.

Shared utility for any resource server’s compute_metrics() override.

Parameters:

tasks
List[List[Dict[str, Any]]]

tasks[i] is a list of rollout dicts for task i.

score_fn
Optional[Any]Defaults to None

Callable(result_dict) -> Dict[str, float|bool] returning named scores. Defaults to lambda r: &#123;"accuracy": r["reward"]&#125;.

answer_key
Optional[str]Defaults to None

Field name for extracted answer (enables majority@k and no_answer). If None, majority@k and no_answer are skipped.

Returns: Dict[str, Any]

Metrics, all_score_dicts, score_names, max_k

nemo_gym.reward_profile.compute_perf_summary(
ng_perf_records: typing.List[typing.Dict[str, typing.Any]],
total_rollouts: int
) -> typing.Optional[typing.Dict[str, typing.Any]]

Aggregate per-rollout ng_perf dicts into the perf_summary block (RFC R3).

total_rollouts is the full rollout count for this batch, including rollouts that never produced ng_perf at all — the denominator for overall_observability_coverage and token_observability_coverage.

Returns None when there are no rollouts at all, or when none of them carried ng_perf (observability off for the whole run, or on but nothing was collected) — perf_summary is absent rather than present with a 0.0 coverage either way; a coverage value, when present, is always > 0. Every other stat is included only when at least one rollout reported the underlying field (a provider that never reports cache usage yields no mean_cached_prompt_tokens, for example).

nemo_gym.reward_profile.compute_subset_metrics(
tasks: typing.List[typing.List[typing.Dict[str, typing.Any]]],
subset_key: str,
score_fn: typing.Optional[typing.Any] = None,
answer_key: typing.Optional[str] = None
) -> typing.Dict[str, typing.Any]

Group tasks by a field and compute pass@k metrics per subset.

Returns flat dict with subset-prefixed keys, e.g. "easy/pass@1/accuracy". Skips the per_sample_aggregate key from each subset’s metrics.

Parameters:

tasks
List[List[Dict[str, Any]]]

tasks[i] is a list of rollout dicts for task i.

subset_key
str

Field name in rollout dicts to group by (e.g. "difficulty").

score_fn
Optional[Any]Defaults to None

Passed through to compute_pass_majority_metrics.

answer_key
Optional[str]Defaults to None

Passed through to compute_pass_majority_metrics.

nemo_gym.reward_profile.coverage_by_agent(
rows: typing.List[typing.Dict[str, typing.Any]],
results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Dict[str, typing.Dict[str, typing.Any]]

Coverage for each agent, computed from that agent’s own records.

A run-wide coverage block copied onto every agent tells each of them how much the run masked, which reads as that agent’s own loss. An agent that masked nothing would carry another agent’s count.

Agents are keyed by name; an agent whose every result was masked still gets an entry, so a caller can keep reporting it after the quality metrics drop it.

nemo_gym.reward_profile.highest_k_metrics(
agent_metrics: typing.Dict[str, typing.Any],
pattern: str,
score_names: typing.Optional[typing.List[str]] = None,
exclude_names: typing.Optional[typing.List[str]] = None
) -> typing.Dict[str, typing.Any]

Select the highest-k entries matching a metric pattern.

Finds all keys matching pattern (with &#123;k&#125; as the k placeholder), determines the highest k value, and returns all entries at that k.

Example::

Get highest-k pass@k for accuracy only

highest_k_metrics(am, “pass@{k}”, score_names=[“accuracy”])

→ {“pass@32/accuracy”: 95.0}

Get highest-k pass@1[avg-of-k] for all scores except no_answer, without stats

highest_k_metrics(am, “pass@1[avg-of-{k}]”, exclude_names=[“no_answer”])

→ {“pass@1[avg-of-32]/accuracy”: 94.5, “pass@1[avg-of-32]/symbolic_accuracy”: 93.2}

Parameters:

agent_metrics
Dict[str, Any]

Full agent metrics dict.

pattern
str

Pattern with &#123;k&#125; placeholder, e.g. "pass@&#123;k&#125;" or "pass@1[avg-of-&#123;k&#125;]".

score_names
Optional[List[str]]Defaults to None

If provided, only return entries whose score name (after the last /) is in this list. Stat suffixes (std_dev, std_err, avg_sample) are always excluded.

exclude_names
Optional[List[str]]Defaults to None

Score names to exclude (e.g. ["no_answer"]). Applied after score_names.

Returns: Dict[str, Any]

Dict of matching metrics at the highest k, e.g. &#123;"pass@32/accuracy": 95.0&#125;.

nemo_gym.reward_profile.select_measured(
rows: typing.List[typing.Dict[str, typing.Any]],
results: typing.List[typing.Dict[str, typing.Any]]
) -> typing.Tuple[typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.List[typing.Dict[str, typing.Any]], typing.Dict[str, typing.Any]]

Narrow a rollout set to what quality metrics may be computed from.

Both the aggregation path and gym eval profile go through here, so the same saved rollouts produce the same quality numbers whichever view you look at.

Three things happen. Masked samples leave the quality set, because their reward is not a measurement of the evaluated system. The flag itself is stripped from what remains: it is a flag, not a measurement, and the profiler would otherwise coerce it to an int and publish mean/mask_sample alongside real metrics. And the input rows of exactly those masked results are dropped with them, so the two stay aligned — a masked rollout is absent from the quality set but was never missing from the collection, and must not be reported as an incomplete run or require allow_partial_rollouts to profile.

Only the masked pairs are removed, never “keep what was scored”: a row whose result is genuinely missing has to survive into the quality set so alignment still reports the collection as partial. This function narrows what is measured; it does not decide whether the collection was complete, and callers that enforce completeness must validate the original rows and results before calling it.

Masked rows are returned rather than discarded: completion and coverage accounting still has to see them.

nemo_gym.reward_profile.MASK_SAMPLE_FIELD = 'mask_sample'
nemo_gym.reward_profile._PERF_SUMMARY_LATENCY_STATS: Tuple[str, ...] = ('p50', 'p90', 'p99', 'mean')
nemo_gym.reward_profile._PERF_SUMMARY_NUM_TURNS_STATS: Tuple[str, ...] = ('mean', 'median', 'std', 'min', 'p90', 'max')
nemo_gym.reward_profile._PERF_SUMMARY_TOKEN_FIELDS: Tuple[str, ...] = ('prompt_tokens', 'cached_prompt_tokens', 'completion_tokens', 'reasoning_tokens...
nemo_gym.reward_profile._PERF_SUMMARY_TOKEN_STATS: Tuple[str, ...] = ('mean', 'median', 'total')
nemo_gym.reward_profile.__getattr__ = moved_attr_getter(__name__, {'reward_profile': 'nemo_gym.cli.eval'})