Switchyard
Switchyard routes each LLM call to the cheapest model that can still do the job. The switchyard_model server puts a Switchyard route behind a NeMo Gym model server, so any benchmark Gym supports can run against a router, or against a fixed baseline, without changing the agent.
By default, Gym hosts Switchyard’s server for you inside the model-server process. There is nothing to install: nemo-switchyard is a dependency of this server, not of Gym itself. Switchyard also runs as a NeMo Relay plugin, a LiteLLM plugin, and an embeddable library; the Switchyard README covers those. This page covers the Gym integration.
Quick start
1. Write routes.toml
routes.toml is a Switchyard deployment file, the same TOML that switchyard-server --config reads. It has three layers: llm_clients (how to reach a provider), targets (one upstream model ID and the client that calls it), and routes (one client-visible model ID and the algorithm that picks between targets). This one defines a fixed baseline and the stage router from the Switchyard README, which picks a tier per turn from the agent’s tool activity:
Each route is a model name an eval can target. The TOML schema lists every key, and the routing overview covers the other route types.
2. Run a benchmark against a route
deployment must be an absolute path. It can also come from the SWITCHYARD_DEPLOYMENT environment variable, in which case the override can be dropped. A bad deployment fails at startup with Switchyard’s validation error rather than as a timeout mid-run. gym env start takes the same flags when you want persistent servers for gym eval run --no-serve.
--model names a route, not a model. It must match a route’s id in routes.toml, or every call fails with model_not_found. Which model actually serves each call is Switchyard’s decision. If the agent sets its own model, the server replaces it with the route, so the eval always measures routing.
Compare routing strategies
Run the same benchmark once per route, each with its own output file and condition_dir. Like deployment, condition_dir is opened by the model-server process, so pass an absolute path:
condition_dir records what each run ran under:
Two things to get right when you compare:
- Join rollouts on
_ng_task_indexand_ng_rollout_indexand drop rows that only one run has. A rollout that failed in one condition must not count in the other. - Add judge tokens to the routed run’s cost. Routes that call an LLM judge, such as
llm_classifier, spend tokens on top of the served model’s. They appear inswitchyard-stats.jsonunderclassifier, not in the rollout’susage.
Example: mean reward and token totals per condition
See which model served each call
Enable model-call capture with ++observability_enabled=true and ++model_call_capture_dir=<absolute path>. Each captured exchange records the response, including the model that served it.
Read the served model from the capture records, not from the rollout’s top-level model. Some agent harnesses overwrite that field with the configured policy model name, which here is the route id.
With capture on, Gym also forwards the rollout-attempt id to Switchyard in the x-switchyard-session-id header. Switchyard uses it as the session id in request logs, OpenTelemetry spans, and session-affinity routing, so proxy-side routing decisions can be joined back to the rollout that produced them. Without capture, the id never reaches this server: Gym warns at startup, or refuses to start if you set forward_session_id: true explicitly.
Run the proxy yourself
Set switchyard_base_url instead of deployment to attach to a proxy you run. Gym then starts nothing of its own. You need this when:
- You run more than one worker. Hosted mode rejects
num_workersabove 1, because each worker process would host its own proxy and split session state and stats. - Several model servers should share one proxy. Routes with session affinity are stateful. Per-server proxies would not route like a single one.
- You need a specific Switchyard build, or the proxy must be reachable from another machine. The hosted proxy binds loopback only.
switchyard_base_url and switchyard_api_key can also come from the SWITCHYARD_BASE_URL and SWITCHYARD_API_KEY environment variables. If both switchyard_base_url and deployment are set, Gym attaches and logs a warning that the deployment is unused.
Comparing runs works the same way, with two differences:
switchyard-stats.jsoncounts from the proxy’s start, not the run’s. Itsscopefield says so. For run-level numbers, give each run its own proxy.- Gym cannot see what build an attached proxy runs, so
nemo_switchyard_versionisnullin the manifest. Record it yourself withproxy_provenance, for example++policy_model.responses_api_models.switchyard_model.proxy_provenance.switchyard_commit=<sha>.
Per-session routing decisions
The hosted proxy cannot write Switchyard’s durable routing log. To get per-session decisions on /v1/routing/session-stats, run the proxy yourself with --routing-log-file and add proxy_x_session_id, the header that log keys on, to session_id_headers: