Reading Results

View as Markdown

AgentEvaluator().run(...) returns an AgentEvalResult and writes nothing. Call result.persist() to store it as a run bundle on disk. The object and the bundle hold the same data — use the object for programmatic follow-up, and the bundle (especially report.html) to inspect or share a run.

The result object

1result = await AgentEvaluator().run(tasks=..., target=...)
AttributeWhat it holds
result.run_idstable identifier for this run (e.g. agent-eval-20260715…)
result.summaryaggregated scores and coverage — see below
result.scoresone entry per (task, trial, metric)
result.trialsone entry per trial
result.tasksthe tasks that were evaluated
result.work_dirthe directory the run worked in, where runtimes wrote trial evidence (None for an in-memory run)

The summary

result.summary (an AgentEvalSummary):

  • summary.scores.scores — the aggregates. Each is named <metric.type>.<output> (and view.<name> for a view), with mean, min, max, and std_dev. This is what the guides print:

    1for aggregate in result.summary.scores.scores:
    2 print(f"{aggregate.name}: {aggregate.mean}")
  • summary.metric_coverage — per metric output, how many trials were total / scored / failed / missing, so you can tell a low mean from low coverage.

  • summary.task_metric_values — per task, the individual trial values behind those means, keyed <metric.type>.<output>. Each record carries the trial_id that produced it and its metric value, so you can answer “which tasks were flaky, and on which trial?” without regrouping result.scores yourself:

    1for task_id, by_output in result.summary.task_metric_values.items():
    2 # .get: keys are per task, so a task scored by a different metric simply has none.
    3 print(task_id, [(a.trial_id, a.value) for a in by_output.get("reward.score", [])])

    A value of None is a trial that died before scoring — a trial that did not pass. A trial whose metric failed is absent entirely, because that leaves it unmeasured rather than unsuccessful. Look up by trial_id rather than by position: the two rules above mean lists for different outputs of one task need not be the same length.

    Values keep the type the metric produced them in — a count stays an int, a flag stays a bool, and a judge’s verdict stays a str. Each record’s value_type says which it is (number, label or missing), which is what tells a real NaN apart from a label that reads "NaN", since strict JSON has no NaN literal and both travel as strings. Before doing arithmetic, project with numeric_metric_values, which drops labels and keeps a dead trial’s None:

    1from nemo_evaluator_sdk.agent_eval.results import numeric_metric_values
    2
    3records = result.summary.task_metric_values["task-47"]["reward.score"]
    4scores = numeric_metric_values(records) # [1.0, None, 0.0]
  • summary.task_outcomes(metric_name=None) — the same data as models rather than nested dicts, sorted by task then metric, each naming its own task_id and metric_name. Pass a "<metric.type>.<output>" to narrow to one metric, which is what a report over a single metric wants:

    1for per_task in result.summary.task_outcomes("reward.score"):
    2 for outcome in per_task.outcomes:
    3 values = numeric_metric_values(outcome.trials)
    4 print(per_task.task_id, outcome.metric_name, values)
    • When you narrow, a task the metric never measured is dropped — it was scored by a different metric, so listing it would invent missing coverage.
    • A task that declared the metric but produced no usable value keeps its entry with an empty trials list, because there the coverage really is missing.
    • Unfiltered, every task is returned.
  • summary.task_count, summary.trial_count, summary.score_count.

Per-metric scores

Each entry in result.scores carries: id, run_id, task_id, trial_id, metric_type, status (e.g. completed / failed), outputs (the metric’s named outputs), diagnostics, and metadata. Use these to drill from an aggregate down to the individual (task, metric) that produced it.

Trials

Each entry in result.trials carries: id, task_id, status (completed / partial / failed), output (the agent’s final answer), evidence (trajectory, final state, logs), and metadata. Trials are the durable, scorer-agnostic record — they can be re-scored offline later.

The run bundle

Call result.persist() and it writes these files (the same data, on disk):

FileContents
run.jsonthe run manifest — run id and a map of the artifact files
summary.jsonthe aggregated summary (means / min / max / std-dev and coverage)
scores.jsonlone row per (task, trial, metric)
trials.jsonlone row per trial — output, evidence, status
tasks.jsonlthe tasks that were evaluated
metadata.jsonrun provenance — labels, target identity, timings, SDK version
report.htmla browsable dashboard — open it in a browser

persist() returns a BundleLocation telling you where the bundle landed:

1from nemo_evaluator_sdk.agent_eval.tasks import AgentEvalRunConfig
2
3result = await AgentEvaluator().run(
4 tasks=..., target=..., config=AgentEvalRunConfig(work_dir="./agent-eval-run"),
5)
6location = result.persist()
7# -> ./agent-eval-run/report.html, summary.json, scores.jsonl, trials.jsonl, ...
8print(location.output_dir, location.dashboard_path)

report.html is the fastest way to eyeball a run or hand it to someone else; the .jsonl files are convenient for loading scores and trials into your own tooling.

persist() defaults to work_dir, which is where the run’s trial evidence already lives — that keeps the bundle self-contained, so it survives being moved or copied. Passing an explicit persist("./elsewhere") is supported, but the bundle’s evidence references still point back at the original directory and only resolve while it exists.

A run with no work_dir and no explicit target raises rather than inventing a directory.

report.html is written unless you pass persist(write_dashboard=False), which emits just the JSON/JSONL artifacts and skips the HTML. The .json and .jsonl files are always written.