nemo_gym.reward_profile
nemo_gym.reward_profile
Module Contents
Classes
Functions
Data
API
Mixin providing compute_metrics/get_key_metrics hooks and the aggregate_metrics endpoint.
Inherited by both SimpleResourcesServer and SimpleResponsesAPIAgent so that benchmark-specific metric logic can live on either server type.
Override to compute custom metrics from all verify responses.
Receives verify responses grouped by task: tasks[i] is a list of rollout dicts for task i. Each dict has at minimum reward, plus any custom fields from the verify response (e.g. symbolic_correct, judgement-gen-base).
Use for metrics that need the full dataset at once:
- Confidence intervals (ArenaMetrics)
- Cross-task statistics (std_dev_across_runs)
- pass@k with proper combinatorial computation
The returned dict is merged into agent_metrics. Default: empty dict (no additional metrics).
Override to select headline metrics for this benchmark.
Default: all mean/* entries from agent_metrics.
Bases: BaseNeMoGymCLIConfig
Aggregate per-repeat estimates (e.g. mean/reward) across repeats, per agent.
Treats each repeat’s stat as one observation and reports the statistics across repeats.
Per-agent, per-rollout-index summary stats across all tasks.
Only produced for agents that have more than one rollout index — agents with a single rollout contribute nothing to a repeat-level comparison and are skipped. If no agent qualifies, returns an empty list.
Return (ci_low, ci_high) t-interval at the given confidence level, or None when n <= 1.
Non-destructively cleans metrics output by RewardProfiler for downstream serialization.
How much of the run the quality metrics above were actually computed from.
Empty when nothing was masked, so a run that reports no masking publishes exactly the keys it published before.
Group verify responses by task index, returning a list of per-task rollout lists.
Split into the samples whose reward is a valid measurement, and the rest.
Add avg_sample_std_dev statistics to an existing metrics dict.
Computes the average of per-task standard deviations across k rollouts — a measure of within-task variance that complements the across-run variance (std_dev_across_runs).
Modifies metrics in place.
Shared aggregation logic for /aggregate_metrics.
RewardProfiler runs with defaults to produce baseline stats (mean/max/min/median/std) for both group-level (per-task) and agent-level metrics.
Compute pass@k, majority@k, no_answer, and variance statistics from grouped task results.
Shared utility for any resource server’s compute_metrics() override.
Parameters:
tasks[i] is a list of rollout dicts for task i.
Callable(result_dict) -> Dict[str, float|bool] returning named scores.
Defaults to lambda r: {"accuracy": r["reward"]}.
Field name for extracted answer (enables majority@k and no_answer). If None, majority@k and no_answer are skipped.
Returns: Dict[str, Any]
Metrics, all_score_dicts, score_names, max_k
Aggregate per-rollout ng_perf dicts into the perf_summary block (RFC R3).
total_rollouts is the full rollout count for this batch, including rollouts that never
produced ng_perf at all — the denominator for overall_observability_coverage and
token_observability_coverage.
Returns None when there are no rollouts at all, or when none of them carried ng_perf
(observability off for the whole run, or on but nothing was collected) — perf_summary is
absent rather than present with a 0.0 coverage either way; a coverage value, when present, is
always > 0. Every other stat is included only when at least one rollout reported the
underlying field (a provider that never reports cache usage yields no
mean_cached_prompt_tokens, for example).
Group tasks by a field and compute pass@k metrics per subset.
Returns flat dict with subset-prefixed keys, e.g. "easy/pass@1/accuracy".
Skips the per_sample_aggregate key from each subset’s metrics.
Parameters:
tasks[i] is a list of rollout dicts for task i.
Field name in rollout dicts to group by (e.g. "difficulty").
Passed through to compute_pass_majority_metrics.
Passed through to compute_pass_majority_metrics.
Coverage for each agent, computed from that agent’s own records.
A run-wide coverage block copied onto every agent tells each of them how much the run masked, which reads as that agent’s own loss. An agent that masked nothing would carry another agent’s count.
Agents are keyed by name; an agent whose every result was masked still gets an entry, so a caller can keep reporting it after the quality metrics drop it.
Select the highest-k entries matching a metric pattern.
Finds all keys matching pattern (with {k} as the k placeholder), determines the
highest k value, and returns all entries at that k.
Example::
Get highest-k pass@k for accuracy only
highest_k_metrics(am, “pass@{k}”, score_names=[“accuracy”])
→ {“pass@32/accuracy”: 95.0}
Get highest-k pass@1[avg-of-k] for all scores except no_answer, without stats
highest_k_metrics(am, “pass@1[avg-of-{k}]”, exclude_names=[“no_answer”])
→ {“pass@1[avg-of-32]/accuracy”: 94.5, “pass@1[avg-of-32]/symbolic_accuracy”: 93.2}
Parameters:
Full agent metrics dict.
Pattern with {k} placeholder, e.g. "pass@{k}" or "pass@1[avg-of-{k}]".
If provided, only return entries whose score name (after the last /)
is in this list. Stat suffixes (std_dev, std_err, avg_sample) are always excluded.
Score names to exclude (e.g. ["no_answer"]). Applied after score_names.
Returns: Dict[str, Any]
Dict of matching metrics at the highest k, e.g. {"pass@32/accuracy": 95.0}.
Narrow a rollout set to what quality metrics may be computed from.
Both the aggregation path and gym eval profile go through here, so the same saved
rollouts produce the same quality numbers whichever view you look at.
Three things happen. Masked samples leave the quality set, because their reward is not
a measurement of the evaluated system. The flag itself is stripped from what remains:
it is a flag, not a measurement, and the profiler would otherwise coerce it to an int
and publish mean/mask_sample alongside real metrics. And the input rows of exactly
those masked results are dropped with them, so the two stay aligned — a masked rollout
is absent from the quality set but was never missing from the collection, and must not
be reported as an incomplete run or require allow_partial_rollouts to profile.
Only the masked pairs are removed, never “keep what was scored”: a row whose result is genuinely missing has to survive into the quality set so alignment still reports the collection as partial. This function narrows what is measured; it does not decide whether the collection was complete, and callers that enforce completeness must validate the original rows and results before calling it.
Masked rows are returned rather than discarded: completion and coverage accounting still has to see them.