Writing Metrics

View as Markdown

A metric scores one trial. Metrics are attached to each task (not once per run), so a suite can grade heterogeneous work — a Q&A task and a coding task can carry different scorers. Every metric, however simple or elaborate, implements the same small protocol.

The protocol

1from nemo_evaluator_sdk.metrics.protocol import MetricInput, MetricOutput, MetricOutputSpec, MetricResult
2
3class MyMetric:
4 @property
5 def type(self) -> str:
6 """A unique name for this metric within a task (also the summary key prefix)."""
7 return "my_metric"
8
9 def output_spec(self) -> list[MetricOutputSpec]:
10 """The named values this metric emits, and their types."""
11 return [MetricOutputSpec.continuous_score("score")]
12
13 async def compute_scores(self, input: MetricInput) -> MetricResult:
14 """Score one trial and return its outputs."""
15 return MetricResult(outputs=[MetricOutput(name="score", value=1.0)])

No base class — a metric is any object with these three members (a structural Metric protocol).

What compute_scores receives

input.candidate is the trial under evaluation:

FieldTypeWhat it holds
candidate.output_textstr | Nonethe agent’s final answer
candidate.evidenceCandidateEvidence | Nonetrajectory, final filesystem state, logs — see Reading evidence
candidate.metadatadicttrial metadata (e.g. a reward a runner stamped on)

input.row.data is a dict describing the task and trial:

KeyWhat it holds
input.row.data["reference"]grader-only ground truth (the task’s reference), never shown to the agent
input.row.data["inputs"]the task inputs (instruction, …)
input.row.data["task"]{id, intent, metadata}
input.row.data["trial"]{id, task_id, status, metadata}

So an outcome metric reads candidate.output_text and row.data["reference"]; a trajectory metric reads candidate.evidence.

Declaring outputs

A metric declares its outputs up front; the runtime validates that compute_scores returns exactly those names, each coercible to the declared type (a missing or undeclared output raises). Build specs with the MetricOutputSpec factories:

FactoryValue typeUse for
MetricOutputSpec.continuous_score(name)floata numeric score (0–1 or unbounded)
MetricOutputSpec.discrete_score(name)intcounts or ordinal levels
MetricOutputSpec.boolean(name)boola pass/fail check
MetricOutputSpec.label(name)stra category label
MetricOutputSpec.model(name, value_schema)your BaseModelstructured/custom values

A metric may emit several outputs — for example an efficiency metric returning both a boolean and a count:

1from nemo_evaluator_sdk.metrics.protocol import MetricOutputSpec
2
3def output_spec(self) -> list[MetricOutputSpec]:
4 return [
5 MetricOutputSpec.boolean("efficient_tool_use"),
6 MetricOutputSpec.discrete_score("max_repeated_tool_calls"),
7 ]

Prefer continuous_score when you want a numeric mean in the run summary. A boolean output reports per-trial pass/fail but does not aggregate to a numeric mean on its own (though it still contributes as 0/1 to a view).

Returning a result

Return a MetricResult whose outputs match output_spec by name:

1from nemo_evaluator_sdk.metrics.protocol import MetricInput, MetricOutput, MetricOutputSpec, MetricResult
2
3class KeywordMatchMetric:
4 @property
5 def type(self) -> str:
6 return "keyword_match"
7
8 def output_spec(self) -> list[MetricOutputSpec]:
9 return [MetricOutputSpec.continuous_score("score")]
10
11 async def compute_scores(self, input: MetricInput) -> MetricResult:
12 expected = str(input.row.data.get("reference", {}).get("expected", "")).lower()
13 answer = (input.candidate.output_text or "").lower()
14 return MetricResult(outputs=[MetricOutput(name="score", value=1.0 if expected and expected in answer else 0.0)])

Reading evidence

candidate.evidence (a CandidateEvidence) holds named descriptors, each exposed through a typed handle. trace() returns the runner’s primary view as an ATIFTraceHandle | OTLPTraceHandle union. Always guard for missing evidence, then narrow on handle.format before using format-specific methods:

1from nemo_evaluator_sdk.metrics.protocol import MetricInput
2from nemo_evaluator_sdk.values.evidence import EVIDENCE_TRACE
3
4async def read_trace(input: MetricInput) -> None:
5 evidence = input.candidate.evidence
6 if evidence is None or evidence.get(EVIDENCE_TRACE) is None:
7 return
8
9 handle = await evidence.trace(EVIDENCE_TRACE)
10 if handle.format == "otlp":
11 resource_spans = await handle.resource_spans()
12 # extract tool calls based on the OTLP exporter specific logic
13 # calls = await extract_tool_calls(resource_spans)
14 else:
15 calls = await handle.tool_calls()
EvidenceHandleReads
trace (EVIDENCE_TRACE)ATIFTraceHandleawait evidence.trace(format="atif").trace(), .tool_calls(), .steps(), .token_usage()
trace (EVIDENCE_TRACE)OTLPTraceHandleawait evidence.trace(format="otlp").resource_spans()
filesystem (EVIDENCE_FINAL_STATE, EVIDENCE_INITIAL_STATE)await evidence.filesystem(name).run_verifier(command), .diff(other) — run a check or diff two snapshots
logsawait evidence.logs(name).read_text(file), .tail(file)

OTLP exposes the unflattened resourceSpans tree; it does not provide a generic tool_calls() method. It’s a responsibility of consumer to parse resourceSpans tree specific to OTLP exporter to extract the searched data.

Asking for one trace format

A runner may record the same trial in more than one encoding. Harbor registers both under trace:atif and trace:otlp when the agent under test emits both. Pass format= to choose one and get a narrowed handle back, with no runtime branch:

1from nemo_evaluator_sdk.values.evidence import CandidateEvidence
2
3async def read_both_views(evidence: CandidateEvidence) -> None:
4 otlp = await evidence.trace(format="otlp") # OTLPTraceHandle
5 resource_spans = await otlp.resource_spans()
6
7 atif = await evidence.trace(format="atif") # ATIFTraceHandle
8 calls = await atif.tool_calls()

trace() with no format still returns whichever encoding the runner made primary. For Harbor that is the OTLP trace when the agent emits one and the ATIF trace otherwise — there is no setting to choose between them. A format= request raises KeyError when the trial carries no trace in that encoding, so a metric that depends on one specific view fails loudly instead of silently scoring nothing.

The Score by Component guide has a complete, runnable trajectory metric; the SDK’s example_metrics.py (under examples/run_agent_eval/) shows filesystem and trace metrics.

How outputs are reported

Each output aggregates in result.summary under the key <metric.type>.<output> (mean / min / max / std-dev). To roll several outputs into one reported score, define a view on the task — see Score by Component.