Reading Results
AgentEvaluator().run(...) returns an AgentEvalResult and writes nothing. Call result.persist() to
store it as a run bundle on disk. The object and the bundle hold the same data — use the object for
programmatic follow-up, and the bundle (especially report.html) to inspect or share a run.
The result object
The summary
result.summary (an AgentEvalSummary):
-
summary.scores.scores— the aggregates. Each is named<metric.type>.<output>(andview.<name>for a view), withmean,min,max, andstd_dev,count, andnan_count.countis the number of finite measured values;nan_countis the number of applicable opportunities that were unmeasured or failed. This is what the guides print: -
summary.metric_coverage— per metric output, how many applicable trials weretotal,scored,missing, orfailed, so you can tell a low mean from low coverage. For each output,scored + missing + failed == total. -
summary.task_metric_values— per task, the individual trial values behind those means, keyed<metric.type>.<output>. Each record carries thetrial_idthat produced it and its metricvalue, so you can answer “which tasks were flaky, and on which trial?” without regroupingresult.scoresyourself:A
valueofNoneis a trial that died before scoring — a trial that did not pass. A trial whose metric failed is absent entirely, because that leaves it unmeasured rather than unsuccessful. An omitted optional output is also absent; it is not stored asNone. Look up bytrial_idrather than by position: these rules mean lists for different outputs of one task need not be the same length.Values keep the type the metric produced them in — a count stays an
int, a flag stays abool, and a judge’s verdict stays astr. Each record’svalue_typesays which it is (number,labelormissing), which is what tells a realNaNapart from a label that reads"NaN", since strict JSON has no NaN literal and both travel as strings. Before doing arithmetic, project withnumeric_metric_values, which drops labels and keeps a dead trial’sNone: -
summary.task_outcomes(metric_name=None)— the same data as models rather than nested dicts, sorted by task then metric, each naming its owntask_idandmetric_name. Pass a"<metric.type>.<output>"to narrow to one metric, which is what a report over a single metric wants:- When you narrow, a task the metric never measured is dropped — it was scored by a different metric, so listing it would invent missing coverage.
- A task that declared the metric but produced no usable value keeps its entry with an empty
trialslist, because there the coverage really is missing. - Unfiltered, every task is returned.
-
summary.task_count,summary.trial_count,summary.score_count.
Sparse output example
Suppose task A has two trials. Its Harbor verifier emits:
reward is measured twice. format_ok is applicable twice but measured once:
These invariants expose the effective denominator:
- For an ordinary aggregate,
count + nan_countequals the applicable opportunities. - If all optional values are omitted, the aggregate, coverage, task-value, view, and pass@k rows remain present and unestimable instead of disappearing.
- Pass@k is calculated separately for every score-like output. Measured
nincludes failed trials as non-passes. A failed metric or an omitted optional output is left out ofn(unmeasured, not unsuccessful). - If task A declares
format_okand task B does not,format_ok’scount,nan_count, and coverage are only over A’s trials. B is omitted from that denominator, not counted as missing.
Per-metric scores
Each entry in result.scores carries: id, run_id, task_id, trial_id, metric_type, status
(e.g. completed / failed), outputs (the metric’s named outputs), diagnostics, and metadata.
Use these to drill from an aggregate down to the individual (task, metric) that produced it.
Trials
Each entry in result.trials carries: id, task_id, status (completed / partial / failed),
output (the agent’s final answer), evidence (trajectory, final state, logs), and metadata. Trials
are the durable, scorer-agnostic record — they can be re-scored offline later.
The run bundle
Call result.persist() and it writes these files (the same data, on disk):
persist() returns a BundleLocation telling you where the bundle landed:
report.html is the fastest way to eyeball a run or hand it to someone else; the .jsonl files are
convenient for loading scores and trials into your own tooling.
An omitted optional output remains absent in scores.jsonl and after loading the bundle. Persistence
does not replace it with JSON null, a NaN label, or 0.0.
persist() defaults to work_dir, which is where the run’s trial evidence already lives — that keeps
the bundle self-contained, so it survives being moved or copied. Passing an explicit
persist("./elsewhere") is supported, but the bundle’s evidence references still point back at the
original directory and only resolve while it exists.
A run with no work_dir and no explicit target raises rather than inventing a directory.
report.html is written unless you pass persist(write_dashboard=False), which emits just the
JSON/JSONL artifacts and skips the HTML. The .json and .jsonl files are always written.
Results from platform jobs
An agent-evaluate platform job persists the same JSON and JSONL data without report.html. It
publishes two named job results:
agent-eval-results— the complete bundle, including tasks, trials, scores, metadata, and summary.summary—summary.jsonby itself for lightweight retrieval.
Wait for the platform job to reach a terminal state before reading either result:
You can also download the bundle through the Jobs CLI:
The queryable agent_eval_results record stores aggregates, coverage, target identity, and the
bundle reference. Individual trials remain in trials.jsonl inside the bundle.