Targets and Runners
AgentEvaluator().run(target=...) accepts one of three kinds of target. Whatever you pick, it produces
trials, and trials are scored the same way — so the same tasks and metrics work against any target
(see Agent Evaluation for the model).
At a glance
The union is AgentEvalTarget = Model | Agent | AgentTaskRunner, where Agent = GenericAgent | NemoAgentToolkitAgent.
Model
A chat/completions endpoint evaluated directly on your tasks — a useful baseline (how well does a
bare model do before you wrap it in an agent?). The evaluator prompts it with each task’s
instruction.
For a local run(), api_key_secret names an environment variable in your process; for a submitted
job it names a platform secret in the workspace.
Agent (HTTP)
A deployed agent reachable over HTTP. Two variants, both authenticated with api_key_secret — the
same credential reference Model uses: for a local run() it names an environment variable, for a
submitted job a platform secret. Its value is sent as a bearer token on each request.
GenericAgent
Any JSON endpoint. You define the request with a Jinja body (rendered against the task inputs) and
pull the answer out with JSONPath. Full walkthrough:
Evaluate a Deployed Agent over HTTP.
NemoAgentToolkitAgent
A NeMo Agent Toolkit endpoint. It
speaks NAT’s fixed request/response protocol, so you don’t hand-write a body — point it at the
workflow’s URL.
AgentTaskRunner (callable, Harbor, Gym, or your own)
The most general target: anything implementing the two-method protocol. Both members are
required — a runner missing either is rejected with NotImplementedError: unsupported agent-eval target type.
The SDK ships three runners you’ll usually reach for first:
CallableAgentTaskRunnerwraps anasync def agent(task) -> str | AgentOutput | TrialDraft. The smallest possible target — no Docker, no HTTP. See the Quickstart. Return aTrialDraftto attach a trajectory or other evidence (see Score by Component).HarborAgentTaskRunnerruns a Harbor task suite in Docker and scores its verifier reward. Pointagent_import_pathat your own Harbor agent, pass its constructor arguments asagent_kwargs(Harbor’s--ak), and hand it credentials throughagent_env_from_host(orenv_secretson the platform’sHarborRunnerTarget). See Harbor Task Suite.GymAgentTaskRunnerruns a NeMo Gym environment and agent, and scores each rollout’s reward. See NeMo Gym Environment.
GymAgentTaskRunner
Runs an existing NeMo Gym environment against your tasks and
adapts its rollouts into trials, scoring each rollout’s reward. GymRuntimeConfig requires agent,
agent_config, and resources_server; discover_gym_tasks builds the tasks from a Gym jsonl
dataset.
Gym is not a dependency of this SDK — it imports Ray at module load, which nemo-platform excludes
by constraint. Install it into its own environment and put its bin on PATH.
It can also be submitted as a platform job from the live runner object, via
client.evaluator.submit(tasks=..., target=runner), rather than described again as a spec.
Platform job specifications use GymRunnerTarget. A submission builds one from two places, split by
what the setting is about:
- The runner carries the evaluation — including
env_secrets, which names environment variables whose values come from a secret reference. That means the same thing wherever the runner runs; only the resolver differs, so a local run resolves it from your environment. - A
GymPlacement, passed tosubmit, carries what the deployment decides:environment(a staged FileSet, requiring sandboxed execution) andagent_ref_name(the agent instance a sandboxed host routes rollouts to).
The overloads accept a GymPlacement only with a GymAgentTaskRunner, so a placement cannot be
paired with a runner that has no use for it.
Full setup, configuration reference, and caveats: Evaluate a NeMo Gym Environment. For custom packages, see Run a Custom Gym Environment.
NeMo Fabric runtimes
NeMo Fabric drives an agent harness rather than a single
agent. Which harness runs is chosen entirely by config["harness"]["adapter_id"], so one runtime
covers several agent frontends:
FabricAgentRuntime runs that config on the host by default, capturing an ATIF trajectory. Pass
sandbox=<SandboxProvider> to run the same configuration inside a sandbox instead; image and
secrets apply there.
It needs the fabric extra, which pulls the Codex, Claude, and Hermes adapters:
The deepagents adapter is deliberately excluded from that extra — it does not support the Relay
observability configuration Fabric streaming generates — so nvidia.fabric.langchain.deepagents
needs its harness installed separately.
Full setup, the agent-config shape, and the trajectory evidence: Evaluate with a NeMo Fabric Harness.
Writing your own
Write your own when your agent doesn’t fit those — a bespoke harness, a queue, a replay of stored runs.
Return one AgentEvalTrial per task and identify the runner with runner_info; the evaluator scores
the trials exactly like any other target:
runner_info is what records the producer of a run: the result carries it on
AgentEvalResult.metadata.target, so a stored run can be understood after the fact. Return a stable
short name ("gym", "harbor") rather than a class name, and keep secrets out of config — it is
persisted with the run bundle.
Choosing a target
- Just trying the flow, or you already have the agent in Python →
CallableAgentTaskRunner. - The agent is deployed behind HTTP →
GenericAgent(any endpoint) orNemoAgentToolkitAgent(a NAT workflow). - You want a model baseline, no agent →
Model. - You have Harbor task datasets →
HarborAgentTaskRunner. - You have a NeMo Gym environment →
GymAgentTaskRunner. - None of the above fits → implement
AgentTaskRunner.