Targets and Runners
AgentEvaluator().run(target=...) accepts one of three kinds of target. Whatever you pick, it produces
trials, and trials are scored the same way — so the same tasks and metrics work against any target
(see Agent Evaluation for the model).
At a glance
The union is AgentEvalTarget = Model | Agent | AgentTaskRunner, where Agent = GenericAgent | NemoAgentToolkitAgent.
Model
A chat/completions endpoint evaluated directly on your tasks — a useful baseline (how well does a
bare model do before you wrap it in an agent?). The evaluator prompts it with each task’s
instruction.
For a local run(), api_key_secret names an environment variable in your process; for a submitted
job it names a platform secret in the workspace.
Agent (HTTP)
A deployed agent reachable over HTTP. Two variants, both authenticated with api_key_secret — the
same credential reference Model uses: for a local run() it names an environment variable, for a
submitted job a platform secret. Its value is sent as a bearer token on each request.
GenericAgent
Any JSON endpoint. You define the request with a Jinja body (rendered against the task inputs) and
pull the answer out with JSONPath. Full walkthrough:
Evaluate a Deployed Agent over HTTP.
NemoAgentToolkitAgent
A NeMo Agent Toolkit endpoint. It
speaks NAT’s fixed request/response protocol, so you don’t hand-write a body — point it at the
workflow’s URL.
AgentTaskRunner (callable, Harbor, Gym, or your own)
The most general target: anything implementing the two-method protocol. Both members are
required — a runner missing either is rejected with NotImplementedError: unsupported agent-eval target type.
The SDK ships three runners you’ll usually reach for first:
CallableAgentTaskRunnerwraps anasync def agent(task) -> str | AgentOutput | TrialDraft. The smallest possible target — no Docker, no HTTP. See the Quickstart. Return aTrialDraftto attach a trajectory or other evidence (see Score by Component).HarborAgentTaskRunnerruns a Harbor task suite in Docker and scores its verifier reward. See Harbor Task Suite.GymAgentTaskRunnerruns a NeMo Gym environment and agent, and scores each rollout’s reward. See NeMo Gym Environment.
GymAgentTaskRunner
Runs an existing NeMo Gym environment against your tasks and
adapts its rollouts into trials, scoring each rollout’s reward. GymRuntimeConfig requires agent,
agent_config, and resources_server; discover_gym_tasks builds the tasks from a Gym jsonl
dataset.
Gym is not a dependency of this SDK — it imports Ray at module load, which nemo-platform excludes
by constraint. Install it into its own environment and put its bin on PATH.
It can also be submitted as a platform job from the live runner object, via
client.evaluator.submit(tasks=..., target=runner), rather than described again as a spec.
Full setup, configuration reference, and caveats: Evaluate a NeMo Gym Environment.
NeMo Fabric runtimes
NeMo Fabric drives an agent harness rather than a single
agent. Which harness runs is chosen entirely by config["harness"]["adapter_id"], so one runtime
covers several agent frontends:
Two runtimes share that config:
FabricAgentRuntimeruns the harness directly, capturing an ATIF trajectory.FabricContainerRuntimeruns the same configuration inside a sandbox, and additionally acceptsskillsandsecrets.
Both need the fabric extra, which pulls the Codex, Claude, and Hermes adapters:
The deepagents adapter is deliberately excluded from that extra — it does not support the Relay
observability configuration Fabric streaming generates — so nvidia.fabric.langchain.deepagents
needs its harness installed separately.
Full setup, the agent-config shape, and the trajectory evidence: Evaluate with a NeMo Fabric Harness.
Writing your own
Write your own when your agent doesn’t fit those — a bespoke harness, a queue, a replay of stored runs.
Return one AgentEvalTrial per task and identify the runner with runner_info; the evaluator scores
the trials exactly like any other target:
runner_info is what records the producer of a run: the result carries it on
AgentEvalResult.metadata.target, so a stored run can be understood after the fact. Return a stable
short name ("gym", "harbor") rather than a class name, and keep secrets out of config — it is
persisted with the run bundle.
Choosing a target
- Just trying the flow, or you already have the agent in Python →
CallableAgentTaskRunner. - The agent is deployed behind HTTP →
GenericAgent(any endpoint) orNemoAgentToolkitAgent(a NAT workflow). - You want a model baseline, no agent →
Model. - You have Harbor task datasets →
HarborAgentTaskRunner. - You have a NeMo Gym environment →
GymAgentTaskRunner. - None of the above fits → implement
AgentTaskRunner.