> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo-platform/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo-platform/_mcp/server.

# Reading Results

> Reference for what a run returns — the in-memory AgentEvalResult (summary, per-metric scores, trials, run id) — and the on-disk run bundle you get by calling persist(), including the browsable HTML report.

`AgentEvaluator().run(...)` returns an `AgentEvalResult` and writes nothing. Call `result.persist()` to
store it as a **run bundle** on disk. The object and the bundle hold the same data — use the object for
programmatic follow-up, and the bundle (especially `report.html`) to inspect or share a run.

## The result object

```python
result = await AgentEvaluator().run(tasks=..., target=...)
```

| Attribute         | What it holds                                                                                      |
| ----------------- | -------------------------------------------------------------------------------------------------- |
| `result.run_id`   | stable identifier for this run (e.g. `agent-eval-20260715…`)                                       |
| `result.summary`  | aggregated scores and coverage — see [below](#the-summary)                                         |
| `result.scores`   | one entry per **(task, trial, metric)**                                                            |
| `result.trials`   | one entry per **trial**                                                                            |
| `result.tasks`    | the tasks that were evaluated                                                                      |
| `result.work_dir` | the directory the run worked in, where runtimes wrote trial evidence (`None` for an in-memory run) |

### The summary

`result.summary` (an `AgentEvalSummary`):

* **`summary.scores.scores`** — the aggregates. Each is named `<metric.type>.<output>` (and
  `view.<name>` for a [view](/documentation/evaluate-models/agent-eval/score-by-component)), with
  `mean`, `min`, `max`, and `std_dev`. This is what the guides print:

  ```python
  for aggregate in result.summary.scores.scores:
      print(f"{aggregate.name}: {aggregate.mean}")
  ```

* **`summary.metric_coverage`** — per metric output, how many trials were `total` / `scored` / failed /
  missing, so you can tell a low mean from low coverage.

* **`summary.task_metric_values`** — per task, the individual trial values behind those means, keyed
  `<metric.type>.<output>`. Each record carries the `trial_id` that produced it and its metric `value`, so
  you can answer "which tasks were flaky, and on which trial?" without regrouping `result.scores`
  yourself:

  ```python
  for task_id, by_output in result.summary.task_metric_values.items():
      # .get: keys are per task, so a task scored by a different metric simply has none.
      print(task_id, [(a.trial_id, a.value) for a in by_output.get("reward.score", [])])
  ```

  A `value` of `None` is a trial that died before scoring — a trial that did not pass. A trial
  whose *metric* failed is absent entirely, because that leaves it unmeasured rather than
  unsuccessful. Look up by `trial_id` rather than by position: the two rules above mean lists for
  different outputs of one task need not be the same length.

  Values keep the type the metric produced them in — a count stays an `int`, a flag stays a `bool`,
  and a judge's verdict stays a `str`. Each record's `value_type` says which it is (`number`,
  `label` or `missing`), which is what tells a real `NaN` apart from a label that reads `"NaN"`, since
  strict JSON has no NaN literal and both travel as strings. Before doing arithmetic, project with
  `numeric_metric_values`, which drops labels and keeps a dead trial's `None`:

  ```python
  from nemo_evaluator_sdk.agent_eval.results import numeric_metric_values

  records = result.summary.task_metric_values["task-47"]["reward.score"]
  scores = numeric_metric_values(records)   # [1.0, None, 0.0]
  ```

* **`summary.task_outcomes(metric_name=None)`** — the same data as models rather than nested dicts,
  sorted by task then metric, each naming its own `task_id` and `metric_name`. Pass a
  `"<metric.type>.<output>"` to narrow to one metric, which is what a report over a single metric
  wants:

  ```python
  for per_task in result.summary.task_outcomes("reward.score"):
      for outcome in per_task.outcomes:
          values = numeric_metric_values(outcome.trials)
          print(per_task.task_id, outcome.metric_name, values)
  ```

  * When you narrow, a task the metric never measured is **dropped** — it was scored by a different
    metric, so listing it would invent missing coverage.
  * A task that declared the metric but produced no usable value keeps its entry with an empty
    `trials` list, because there the coverage really is missing.
  * Unfiltered, every task is returned.

* **`summary.task_count`**, **`summary.trial_count`**, **`summary.score_count`**.

### Per-metric scores

Each entry in `result.scores` carries: `id`, `run_id`, `task_id`, `trial_id`, `metric_type`, `status`
(e.g. `completed` / `failed`), `outputs` (the metric's named outputs), `diagnostics`, and `metadata`.
Use these to drill from an aggregate down to the individual (task, metric) that produced it.

### Trials

Each entry in `result.trials` carries: `id`, `task_id`, `status` (`completed` / `partial` / `failed`),
`output` (the agent's final answer), `evidence` (trajectory, final state, logs), and `metadata`. Trials
are the durable, scorer-agnostic record — they can be re-scored offline later.

## The run bundle

Call `result.persist()` and it writes these files (the same data, on disk):

| File            | Contents                                                          |
| --------------- | ----------------------------------------------------------------- |
| `run.json`      | the run manifest — run id and a map of the artifact files         |
| `summary.json`  | the aggregated summary (means / min / max / std-dev and coverage) |
| `scores.jsonl`  | one row per (task, trial, metric)                                 |
| `trials.jsonl`  | one row per trial — output, evidence, status                      |
| `tasks.jsonl`   | the tasks that were evaluated                                     |
| `metadata.json` | run provenance — labels, target identity, timings, SDK version    |
| `report.html`   | a browsable dashboard — open it in a browser                      |

`persist()` returns a `BundleLocation` telling you where the bundle landed:

```python
from nemo_evaluator_sdk.agent_eval.tasks import AgentEvalRunConfig

result = await AgentEvaluator().run(
    tasks=..., target=..., config=AgentEvalRunConfig(work_dir="./agent-eval-run"),
)
location = result.persist()
# -> ./agent-eval-run/report.html, summary.json, scores.jsonl, trials.jsonl, ...
print(location.output_dir, location.dashboard_path)
```

`report.html` is the fastest way to eyeball a run or hand it to someone else; the `.jsonl` files are
convenient for loading scores and trials into your own tooling.

`persist()` defaults to `work_dir`, which is where the run's trial evidence already lives — that keeps
the bundle self-contained, so it survives being moved or copied. Passing an explicit
`persist("./elsewhere")` is supported, but the bundle's evidence references still point back at the
original directory and only resolve while it exists.

A run with no `work_dir` and no explicit target raises rather than inventing a directory.

`report.html` is written unless you pass `persist(write_dashboard=False)`, which emits just the
JSON/JSONL artifacts and skips the HTML. The `.json` and `.jsonl` files are always written.

## Related

#### [Quickstart](/documentation/evaluate-models/agent-eval/quickstart)

#### [Score by Component](/documentation/evaluate-models/agent-eval/score-by-component)

#### [Writing Metrics](/documentation/evaluate-models/agent-eval/writing-metrics)