Configuration

View as Markdown

Complete syntax and field specifications for NeMo Gym configuration files.

File Locations

FileLocationVersion Control
Server configs<server_type>/<implementation>/configs/*.yaml✅ Committed
env.yamlRepository root (./env.yaml)❌ Gitignored (user creates)

Artifact Roots (results_dir, cache_dir)

Two opt-in top-level config keys that servers can consult for state that outlives a request:

KeyDefaultRead today by
results_dir<working dir>/resultsthe swe_agents server (per-run results) and the W&B run directory
cache_dir<working dir>/cachethe swe_agents server (setup trees) and the uv_cache_dir default

These are opt-in settings, not a global redirect: components that don’t consult them (profiling output, rollout capture, logs, Hugging Face caches, other agents) keep their own locations, and the module-level RESULTS_DIR/CACHE_DIR constants remain unchanged defaults for such consumers. New artifact producers are encouraged to read these keys.

Both are independently overridable (results_dir=/shared/results cache_dir=/shared/cache), and relative values are normalized to absolute paths at parse time so all server processes agree. uv_venv_dir keeps its own key and default.

On multi-node deployments, the swe_agents server embeds these paths in commands executed from other nodes, so both roots must resolve to the same content at the same path on every node that runs rollout workers: use a shared filesystem, or pre-stage identical trees at the same path on each node (for example, baked into the container image). A cache that exists only on the head node breaks remote rollouts.

Servers that pre-staged setup trees next to the package (the layout before cache_dir existed, for example baked container images) keep using those trees only while cache_dir is left at its default — explicitly configuring cache_dir opts out, and removing the pre-staged trees migrates a default-config deployment. Contents accumulate across runs; both directories are safe to delete when no run is active.


Path Resolution and External Roots

Gym resolves every relative path — config_paths, env.yaml, prompt configs, dataset files, the --<component> selectors (--benchmark, --environment, --model-type, --resources-server), and server directories (used by gym env test) — against an ordered list of roots, returning the first one where the path exists:

  1. Extra roots from NEMO_GYM_EXTRA_ROOTS (or --search-dir), in the order listed.
  2. The current working directory — your project.
  3. The Gym install root, where the built-in components live (in both editable and wheel installs).

Earlier roots win, so a component you provide shadows a same-named built-in. Absolute paths are used unchanged.

External roots (NEMO_GYM_EXTRA_ROOTS)

Point Gym at one or more extra roots so your own benchmarks, environments, resources servers, agents, models, configs, prompts, and data resolve by name — without forking Gym or copying files into the install tree. Each root uses the same layout as the Gym repo:

<root>/
├── benchmarks/<name>/
├── environments/<name>/
├── resources_servers/<name>/
├── responses_api_agents/<name>/
└── responses_api_models/<name>/

Set it as an os.pathsep-separated list (: on Linux/macOS):

export NEMO_GYM_EXTRA_ROOTS=~/my-plugins:~/team-benchmarks
gym list benchmarks # your benchmarks appear alongside the built-ins
gym eval run --benchmark my_bench --model-type vllm_model

The variable is inherited by the servers Gym spawns, so plugin components resolve inside them too.

--search-dir

--search-dir DIR (repeatable) is the per-invocation equivalent: Gym sets NEMO_GYM_EXTRA_ROOTS to its value for the duration of that command, then restores it. Use it for one-off runs instead of exporting the variable.

gym list benchmarks --search-dir ~/my-plugins
gym eval run --benchmark my_bench --model-type vllm_model --search-dir ~/my-plugins

Prefer NEMO_GYM_EXTRA_ROOTS when the same plugin roots apply to every command in a shell session; reach for --search-dir for a single invocation. If both are set, --search-dir takes precedence for that command.


Server Configuration

All servers share this structure:

server_id: # Your unique name for this server
server_type: # responses_api_models | resources_servers | responses_api_agents
implementation: # Directory name inside the server type directory
entrypoint: app.py # Python file to run
# ... additional fields vary by server type

Reverse Proxy Headers

Gym disables Uvicorn proxy-header processing by default. Direct Gym traffic does not pass through a reverse proxy, so the network peer address and request scheme are authoritative without interpreting X-Forwarded-For or X-Forwarded-Proto.

Deployments behind a reverse proxy must opt in with top-level configuration and list the proxy addresses or CIDRs that are allowed to supply forwarded headers:

uvicorn_proxy_headers: true
uvicorn_forwarded_allow_ips:
- 10.0.0.5
- 10.20.0.0/24

Both settings apply to the Gym head server and the Agent, Model, and Resources servers. Enabling proxy headers without a non-empty allowlist fails at startup. Wildcards and all-address CIDRs such as 0.0.0.0/0 and ::/0 are rejected.

Only list reverse proxies you control. Trusting a loopback address such as 127.0.0.1 trusts every local process, not only a proxy running on that address. Deployments that previously relied on Uvicorn’s implicit loopback trust must set these options explicitly.

Model Server Fields

policy_model: # Server ID (use "policy_model" — agent configs expect this name)
responses_api_models: # Server type (must be "responses_api_models" for model servers)
openai_model: # Implementation (use "openai_model", "vllm_model", or "azure_openai_model")
entrypoint: app.py # Python file to run
openai_base_url: ${policy_base_url} # API endpoint URL
openai_api_key: ${policy_api_key} # Authentication key
openai_model: ${policy_model_name} # Model identifier

Keep the server ID as policy_model — agent configs reference this name by default. The ${policy_base_url}, ${policy_api_key}, and ${policy_model_name} placeholders should be defined in env.yaml at the repository root, allowing you to change model settings in one place.

Resources Server Fields

my_resource: # Server ID (your choice — agents reference this name)
resources_servers: # Server type (must be "resources_servers" for resources servers)
example_single_tool_call: # Implementation (must match a directory in resources_servers/)
entrypoint: app.py # Python file to run
domain: agent # Server category (see values below)
verified: false # Passed reward profiling and training checks (default: false)
description: "Short description" # Server description
value: "What this improves" # Training value provided
allowed_agents: [my_agent] # Optional: agents that score this task correctly

allowed_agents pins the agent harnesses that verify this server’s tasks correctly. Omit it (the default) when any harness will do. When it is set, --agent-type refuses to substitute an agent that is not listed. Declare it on every config that writes this server’s block — server configs are self-contained, so a config written from scratch does not inherit it.

Domain values: math, coding, agent, knowledge, instruction_following, long_context, safety, games, translation, e2e, rlhf, other (see Domain)

Agent Server Fields

Agent servers must include both a resources_server and model_server block to specify which servers to use.

my_agent: # Server ID (your choice — used in API requests)
responses_api_agents: # Server type (must be "responses_api_agents" for agent servers)
simple_agent: # Implementation (must match a directory in responses_api_agents/)
entrypoint: app.py # Python file to run
resources_server: # Specifies which resources server to use
type: resources_servers # Always "resources_servers"
name: my_resource # Server ID of the resources server
model_server: # Specifies which model server to use
type: responses_api_models # Always "responses_api_models"
name: policy_model # Server ID of the model server
datasets: # Optional: define for training workflows
- name: train # Dataset identifier
type: train # example | train | validation
jsonl_fpath: path/to/data.jsonl # Path to data file
license: Apache 2.0 # Required for train/validation

Dataset Configuration

Define datasets associated with agent servers for training and evaluation.

datasets:
- name: my_dataset
type: train
jsonl_fpath: path/to/data.jsonl
license: Apache 2.0
num_repeats: 1
FieldRequiredDescription
nameYesDataset identifier
typeYesexample, train, or validation
jsonl_fpathYesPath to data file
licenseFor train/validationLicense identifier (see values below)
sourceNoWhere to fetch the data from when it’s missing locally. A source: block with type: gitlab (dataset_name, version, artifact_fpath) or type: huggingface (repo_id, optional artifact_fpath). Replaces the deprecated gitlab_identifier: / huggingface_identifier: blocks.
num_repeatsNoRepeat dataset n times (default: 1)

Dataset types:

  • example — For testing and development
  • train — Training data (requires license)
  • validation — Evaluation data (requires license)

License values: Apache 2.0, MIT, Creative Commons Attribution 4.0 International, Creative Commons Attribution-ShareAlike 4.0 International, CC BY-SA 4.0, CC BY-NC 3.0, TBD (see license)


Local Configuration (env.yaml)

Store secrets and local settings at the repository root. This file is gitignored.

# Policy model (required for most setups)
# Reference these variables in server configs using `${variable_name}` syntax (e.g., `${policy_base_url}`)
policy_base_url: https://api.openai.com/v1
policy_api_key: sk-your-api-key
policy_model_name: gpt-4o-2024-11-20
# Optional: store config paths for reuse
my_config_paths:
- responses_api_models/openai_model/configs/openai_model.yaml
- resources_servers/example_single_tool_call/configs/example_single_tool_call.yaml
# Optional: multi-node setups
use_absolute_ip: true # Bind servers to the host's IP instead of 127.0.0.1 (default: false)
# Optional: validation behavior
error_on_almost_servers: true # Exit on invalid configs (default: true)
# Optional: experiment tracking (see Experiment Tracking below)
wandb_project: gym-dev
wandb_name: my-run
wandb_api_key: your-wandb-key

Multi-Node Configuration

use_absolute_ip — Controls the default host servers bind to.

  • Default: false — servers use 127.0.0.1 (localhost).
  • When to use: Set to true for multi-node setups (e.g. multi-node Ray clusters) where servers must communicate across machines.
  • Effect: Resolves and uses the host’s IP address (gethostbyname(gethostname())) instead of localhost.

Experiment Tracking

Gym exports the resolved config, aggregate metrics, and rollouts to any tracking backend you configure. Weights & Biases and MLflow are both supported, and both can be on at once. A backend is used only when every key it needs is set; otherwise it is silently skipped.

Weights & Biases

KeyRequiredDescription
wandb_projectyesW&B project to log to.
wandb_nameyesRun name.
wandb_api_keyyesW&B API key.

MLflow

KeyRequiredDescription
mlflow_tracking_uriyesTracking server URI.
mlflow_experiment_nameyesExperiment to log to. Created if it does not exist.
mlflow_run_nameyesRun name.
mlflow_tracking_tokennoBearer token. Omit for unauthenticated servers.
gym eval run \
--benchmark gpqa \
--model-type openai_model \
+mlflow_tracking_uri=https://mlflow.example.com/ \
+mlflow_experiment_name=my-experiment \
+mlflow_run_name=gpqa

mlflow_tracking_uri and mlflow_tracking_token are shared with the GitLab model registry used by gym dataset upload|download. Setting only those two does not enable the exporter — it also needs an experiment and run name.

Rollout upload

Rollouts are uploaded to every configured backend by default. Turn this off when they are large:

gym eval run ... +upload_rollouts=false

Metrics and config are still exported. The old name upload_rollouts_to_wandb is deprecated and will be removed in a future release.

Exporting is best-effort. Errors that happen during export (e.g., unavailable server) are reported as warnings and the run continues.


Command Line Usage

To run servers, use gym env start. NeMo Gym uses Hydra for configuration management.

Loading Configs

# Load one or more config files
gym env start \
--config config1.yaml \
--config config2.yaml
# Use paths stored in env.yaml
gym env start "+config_paths=${my_config_paths}"

Overriding Values

# Override nested values (use dot notation after server ID)
gym env start \
--config config.yaml \
+my_server.resources_servers.my_impl.domain=coding
# Override policy model
gym env start --config config.yaml \
--model gpt-4o-mini
# Disable strict validation
gym env start \
--config config.yaml \
+error_on_almost_servers=false

Troubleshooting

Configuration for common configuration errors and solutions.