CLI Commands

View as Markdown

This page documents the NeMo Gym command-line interface.

All functionality is exposed through a single gym entry point, organized into command groups (gym <group> <command>). ng is a drop-in alias for gym. Every group and command supports -h/--help.

The legacy ng_* / nemo_gym_* commands (such as ng_run or nemo_gym_collect_rollouts) still work but are deprecated. Each one prints a notice pointing at its gym replacement and then runs it. See Migrating from the legacy commands at the bottom of this page.

Quick Reference

# General
gym --help # list all command groups
gym --version [--json] # print version and system info
ng ... # 'ng' is an alias for 'gym'
# Discover
gym list benchmarks [<name>] [--json] # list benchmarks (or inspect one), by their --benchmark value
gym list environments [<name>] [--json] # list the unified environment and benchmark catalog
gym list agents [<name>] [--json] # list agent harnesses and how each composes
gym list models [<name>] [--json] # list --model-type values (or inspect a model)
gym list resources-servers [<name>] [--json] # list --resources-server values (or inspect one)
gym search [<type>] <query> [--json] # search the unified environment catalog, or one component type
# Datasets
gym dataset upload # upload a prepared dataset to HF (default) or GitLab
gym dataset download # download a dataset from HF (default) or GitLab
gym dataset rm # delete a dataset from GitLab
gym dataset migrate # move a dataset from GitLab to HF
gym dataset render # generate a dataset preview (materialize prompts)
gym dataset collate # validate and collate a dataset
# Environments
gym env init # scaffold a resources server, environment, or benchmark
gym env resolve # resolve and print the final merged config
gym env validate # validate a manifest-backed workload or legacy config without services
gym env packages # list packages in a server's virtual environment
gym env test # run resources-server tests
gym env publish # finalize readiness and confirm registry discovery
gym env start # start the servers
gym env status # show running servers
# Evaluation
gym eval prepare # prepare benchmark data and dump it to disk
gym eval run # collate data, start servers, and collect rollouts
gym eval aggregate # merge sharded rollout results
gym eval export # export supported Gym trajectories as ATIF
gym eval health-check # verify rollout artifact quality for an existing run
gym eval profile # compute a reward profile from rollouts
gym eval reverify # recompute rewards from existing rollouts without re-running inference
gym eval compare # compare a baseline eval run against a candidate run
# Contributor helpers
gym dev test # run NeMo Gym's unit tests

Common Options

These options are shared across many commands.

OptionDescription
--config PATHLoad a Gym config YAML. Repeatable. Maps to +config_paths=[...].
--benchmark NAMESelect a registered benchmark by name instead of a config path.
--resources-server NAMESelect a registered resources server by name.
--model-type NAMESelect a registered model server type by name (such as openai_model or vllm_model).
--agent-type NAME[/FLAVOR]Run the environment with a different agent harness. Refer to Swapping the agent.
--search-dir DIRExtra root directory to search for components and configs. Repeatable. Applies everywhere Gym resolves paths — discovery, the --<component> selectors, config_paths, prompt configs, and data files — and is inherited by spawned servers. Equivalent to the NEMO_GYM_EXTRA_ROOTS environment variable; see Configuration.
--jsonEmit machine-readable JSON instead of human-readable output (reporting commands only). Accepted before or after the subcommand; commands that emit no JSON reject it.
-v, --verboseSet the logging level to DEBUG. Flows through to spun-up servers.
-h, --helpShow help for any group or command.

The --benchmark, --resources-server, --model-type, and --agent-type selectors resolve a component name to its config file for you, so you do not need to know the project’s directory layout. If you mistype a name, the CLI suggests the closest match. To point at a config file directly, use --config <path> instead.

Selecting the model server

Commands that need a model (gym env start, gym eval run) configure it through four flags:

FlagDescription
--model-type NAMEThe model server type to load (such as openai_model, vllm_model, or local_vllm_model).
--model, -mThe served model identifier: an API model name, an HF id, or a local checkpoint path. Interpreted per --model-type. Maps to policy_model_name.
--model-urlBase URL of an existing model server endpoint. Maps to policy_base_url.
--model-api-keyAPI key for the model server. Maps to policy_api_key.

Swapping the agent

--agent-type NAME runs an environment or benchmark with a different agent harness, without editing any config. It resolves to responses_api_agents/NAME/configs/NAME.yaml; use NAME/FLAVOR to pick a non-default config in that directory.

# Run Terminal-Bench 2.1 under Terminus 2 instead of the OpenCode harness its config names
gym eval run \
--benchmark terminal_bench_2_1/opencode \
--agent-type terminus_2_sandboxed_agent \
--model-type vllm_model

Available on gym env start, gym env validate, gym env prefetch, and gym eval run. Like the other selectors it is pure shorthand, so --agent-type terminus_2_sandboxed_agent and --config <that file> behave identically.

Run gym list agents to see the harnesses you can pass.

The composed server instance is renamed after the agent that runs it, so metrics and gym env status report the harness actually used: composing terminus_2_sandboxed_agent onto terminal_bench_2_1_opencode_sandboxed_agent produces terminal_bench_2_1_terminus_2_sandboxed_agent. gym env start prints the instance list, which matters for the two-step flow:

gym env start \
--benchmark terminal_bench_2_1/opencode \
--agent-type terminus_2_sandboxed_agent \
--model-type vllm_model
gym eval run --no-serve --agent terminal_bench_2_1_terminus_2_sandboxed_agent --input rows.jsonl

--agent-type chooses the harness to compose; --agent names the running instance to collect rollouts from. Because the harness is settled once the servers are up, --agent-type is rejected with --no-serve.

Anything that routes rows — --agent, +agent_map values, +fan_out entries — must name the composed instance, and Gym says so rather than failing later:

Routing names agent instances that no longer exist once the agent is swapped:
- agent_name: 'terminal_bench_2_1_opencode_sandboxed_agent' is now
'terminal_bench_2_1_terminus_2_sandboxed_agent'
Use the name the composed config reports.

Rows collected before the swap are re-routed for you, so existing rollout files stay usable.

Some tasks only work with the harness they were built for. A resources server may pin the harnesses required for the task with allowed_agents, and Gym refuses a pairing it does not list, before any server starts:

'simple_agent' is not declared compatible with 1 of the agent instance(s) it would replace,
so it cannot be scored correctly:
- terminal_bench_2_1_opencode_sandboxed_agent uses terminal_bench_2_1_opencode_resources_server
and accepts opencode_sandboxed_agent, terminus_2_sandboxed_agent
Select one of: opencode_sandboxed_agent, terminus_2_sandboxed_agent. Or pass
--allow-unsupported-pairing (or set NEMO_GYM_ALLOW_UNSUPPORTED_PAIRING=1) to bypass the check.

Results from an undeclared pairing might not be valid.

Hydra overrides (escape hatch)

The gym CLI is a thin wrapper over Gym’s Hydra config system. Standard flags (--flag value) cover the common inputs. Anything not covered by a flag can still be passed as a raw Hydra override using +key=value (add a new key) or ++key=value (add or override an existing key). Unknown overrides are forwarded to Hydra untouched, so advanced config composition is fully functional.

When a command needs overrides for keys that have no dedicated flag, you can keep the whole command in Hydra form (config paths included) or mix --flag and +key=value styles in one invocation:

gym eval run \
--benchmark aime24 \
--model-type openai_model \
++responses_create_params.reasoning.effort=low \
+wandb_project=gym-dev

General

gym --help

List all command groups and their descriptions.

gym --help

gym --version

Print the NeMo Gym version along with Python, key dependency, and system information.

OptionDescription
--jsonOutput version information as JSON.
gym --version
# Output as JSON
gym --version --json

Discovery

Commands for discovering benchmarks, environments, model servers, agent harnesses, and resources servers.

gym list <type> lists all components of that type; gym list <type> <name> inspects that one. All commands accept:

  • --json which prints machine-readable output instead of a human-readable table.
  • --search-dir which allows you to surface your own Gym-compatible components alongside the built-ins.

gym list benchmarks

List the benchmarks available in NeMo Gym, with their kind, manifest status, domain, and description. The value in the Name column is the benchmark identifier and can be passed to --benchmark. To see which agent harness a benchmark composes with, inspect it with gym list benchmarks <name>.

OptionDescription
[<name>]Inspect a single benchmark (config path, manifest status, agent, datasets) instead of listing all.
--jsonOutput the benchmark list as JSON.
gym list benchmarks
gym list benchmarks aime24 # inspect one benchmark
# Machine-readable output for scripting
gym list benchmarks --json | jq '.[].name'

gym list environments

List the manifest-first catalog unioned with legacy environment and benchmark configs. Manifest-backed entries expose their declared metadata; unmigrated entries remain visible with no-manifest status.

OptionDescription
[<name>]Inspect one environment or benchmark.
--kind environment|benchmarkFilter the list or resolve an ambiguous name.
--domain, --modality, --licensingFilter by catalog metadata.
--status experimental|no-manifestFilter by the experimental annotation or missing manifest.
--lifecycle active|deprecatedFilter by lifecycle.
--jsonOutput catalog entries as JSON.
gym list environments
gym list environments calendar # inspect one environment
gym list environments --kind benchmark --domain math
gym list environments --json | jq '.[].name'

gym list agents

List the agent harnesses under responses_api_agents/, with each one’s composition pattern — composable (Pattern A: references a separate resources server, so it can be wired into a matching environment) vs self-contained (Pattern B: ships its own framework/environment) — plus its config variants.

OptionDescription
[<name>]Inspect a single agent (composition pattern and config variants).
--jsonOutput the agent list as JSON (includes the self_contained flag).
gym list agents
gym list agents simple_agent # inspect one agent
gym list agents --json | jq '.[] | {name, self_contained}'

gym list models

List the model servers under responses_api_models/. The value in the Model column corresponds to a config flavor and can be passed as the --model-type flag to other commands.

OptionDescription
[<name>]Inspect a single model instead of listing all.
--jsonOutput the rows as JSON.
gym list models
gym list models vllm_model # inspect one model
gym list models --json | jq '.[].model'

gym list resources-servers

List the resources servers under resources_servers/. The value in the Name column corresponds to the config flavor and can be passed as --resources-server.

OptionDescription
[<name>]Inspect a single resources server (config path, domain, description).
--jsonOutput the list as JSON.
gym list resources-servers
gym list resources-servers mcqa # inspect one resources server
gym list resources-servers --json | jq '.[].name'

Filter a listing to fuzzy matches on a query (against each component’s name/value, description, domain, and type-specific fields). Without a type, the command searches the unified environment and benchmark catalog.

OptionDescription
[<type>]Optional component type: benchmarks, environments, agents, models, or resources-servers.
QUERYText matched (substring or fuzzy) against a component’s name, description, and key metadata (positional, required).
--jsonOutput matches as JSON.
gym search math # environments and benchmarks matching "math"
gym search benchmarks math # only benchmarks matching "math"
gym search environments calendar # environments matching "calendar"
gym search models vllm --json

Datasets

Commands for preparing, previewing, and managing datasets. By default dataset transfer commands use HuggingFace; pass --storage gitlab to target the GitLab Registry.

gym dataset upload

Upload a prepared local JSONL dataset to HuggingFace (default) or GitLab.

OptionDescription
--storage {hf,gitlab}Storage backend. Default: hf.
--input, -iLocal JSONL file to upload.
--nameDataset name.
--revisionDataset revision (version).
--splitDataset split (HF only).
--create-prOpen a pull request with your changes (HF only).
# Upload to HuggingFace
gym dataset upload \
--name my_dataset \
--input data/train.jsonl \
--revision 0.0.1
# Upload to GitLab
gym dataset upload \
--storage gitlab \
--name my_dataset \
--input data/train.jsonl \
--revision 0.0.1

gym dataset download

Download a dataset from HuggingFace (default) or GitLab.

OptionDescription
--storage {hf,gitlab}Storage backend. Default: hf.
--repo-idHF repo id, such as org/dataset (HF only).
--nameDataset name (GitLab only).
--revisionDataset version (GitLab only).
--artifactRemote file to fetch (GitLab: required; HF: optional raw file).
--output, -oLocal destination file.
--output-dirLocal destination directory; needed when downloading all splits (HF only).
--splitDataset split (HF only).
# Download a single file from HuggingFace
gym dataset download --repo-id NVIDIA/NeMo-Gym-Math-example_multi_step-v1 \
--artifact train.jsonl \
--output data/train.jsonl
# Download from GitLab
gym dataset download \
--storage gitlab \
--name example_multi_step \
--revision 0.0.1 \
--artifact train.jsonl \
--output data/train.jsonl

gym dataset rm

Delete a dataset from the GitLab Registry. Prompts for confirmation.

OptionDescription
--nameName of the dataset to delete.
gym dataset rm --name old_dataset

gym dataset migrate

Migrate a JSONL dataset to HuggingFace from GitLab. Use gym dataset upload if you do not want automatic GitLab deletion.

OptionDescription
--input, -iLocal JSONL file to upload to HF.
--nameDataset name.
--revisionDataset revision (HF).
--splitDataset split.
--create-prOpen a pull request to HF dataset with your changes.
gym dataset migrate \
--name my_dataset \
--input data/train.jsonl \
--revision 0.0.1

gym dataset collate

Validate and collate a dataset, generating metrics and statistics.

OptionDescription
--config PATHConfig file to load. Repeatable.
--resources-server NAMELoad the named resources server config.
--search-dir DIRExtra root directory to search for named components. Repeatable.
--mode {train_preparation,example_validation}Use train_preparation to prepare train/validation datasets, or example_validation to validate example data.
--output-dirOutput directory for the prepared data.
--downloadDownload source datasets before collating.
gym dataset collate \
--resources-server example_multi_step \
--output-dir data/example_multi_step \
--mode example_validation

gym dataset render

Generate a dataset preview by materializing prompts from a raw input file and a prompt template, producing JSONL with populated responses_create_params.input for RL training.

Each input row must not already have a populated responses_create_params.input; the command applies the prompt template from --prompt-config to each row, fills in the input, and preserves the row’s other fields.

OptionDescription
--input, -iRaw input JSONL file (rows without responses_create_params.input).
--prompt-configPrompt template YAML to apply.
--output, -oOutput JSONL file.
--search-dir DIRExtra root directory to search for named components. Repeatable.
gym dataset render \
--input raw.jsonl \
--prompt-config prompt.yaml \
--output preview.jsonl

Which data-preparation command should I use?

  • gym dataset render — a focused, standalone step that applies a prompt template to raw rows to populate responses_create_params.input. No servers are started. Use it when you have raw data and just need to turn it into prompt-ready rows.
  • gym dataset collate — the full preparation pipeline for training: it can download missing datasets, validate data, and compute dataset metrics, writing train/validation splits and metrics artifacts. Use it to prepare and validate datasets for training or PR submission.

Environments

Commands for developing, running, and inspecting environments. An environment defines tasks, verification, and its interaction surface; Gym config composes it with agent, model, and runtime components.

gym env init

Scaffold a resources server, or create a manifest-backed environment or benchmark skeleton. Workload scaffolds include a manifest, Gym config, sample data, and the extension points selected by the integration profile.

OptionDescription
--resources-server NAMEName of the resources server to create.
--environment NAME / --benchmark NAMECreate a manifest-backed environment or benchmark.
--profile PROFILEIntegration profile: custom-gym-verifier, custom-gym-agent-loop, external-agent-loop, or external-rollout-driver.
--reuse-verifier NAMEReuse an existing resources server that exports VERIFIER_FIXTURE.
--reward-range LOW HIGHReward endpoints required when reusing a verifier.
--higher-is-better / --lower-is-betterReward direction required when reusing a verifier.
gym env init --resources-server my_server
gym env init --environment my_eval --profile custom-gym-verifier
gym env init --benchmark my_benchmark --reuse-verifier existing_scorer \
--reward-range 0 1 --higher-is-better

gym env resolve

Resolve the configs, flags, and overrides into a final merged config and print it. Useful for debugging configuration. Secrets are hidden.

Unlike gym env validate, resolve substitutes no dummy model, so a model config’s policy_* values must be supplied for its interpolations to resolve.

OptionDescription
--config PATHConfig file to load. Repeatable.
--search-dir DIRExtra root directory to search for named components. Repeatable.
# Merge a resources-server config with a model config, apply an override, and print the result
gym env resolve \
--config resources_servers/example_single_tool_call/configs/example_single_tool_call.yaml \
--config responses_api_models/openai_model/configs/openai_model.yaml \
++responses_create_params.temperature=0.6 \
++policy_base_url=https://api.openai.com/v1 \
++policy_api_key=sk-example \
++policy_model_name=gpt-4o-mini

gym env validate

Validate a manifest-backed workload or legacy config without starting Ray or server subprocesses. Manifest validation resolves the Gym config offline, checks composition mirrors and datasets, and reports the resolved components. Legacy validation retains the existing config pre-flight behavior. Exits 0 when valid, or 1 with a clean message when not.

By default, values a run needs but the config leaves undefined (such as policy_base_url from env.yaml) are filled with placeholders so the rest can still be checked, and reported as a warning. Pass --strict to fail on them instead.

OptionDescription
NAMEManifest-backed catalog entry to validate.
--kind environment|benchmarkResolve an ambiguous catalog name.
--manifest PATHValidate a manifest directly instead of selecting by name.
--syncAtomically synchronize config-owned composition mirrors after all checks pass.
--strictResolve the config exactly as a run would, with no placeholder model values. An undefined value fails instead of warning.
--jsonOutput the manifest validation report as JSON.
--config PATHConfig file to load. Repeatable.
--environment NAME / --benchmark NAMEValidate a named environment / benchmark config.
--resources-server NAMEValidate a named resources server config.
--model-type NAMEAlso load a named model config (otherwise a dummy policy_model is used).
--search-dir DIRExtra root directory to search for named components. Repeatable.
--agent-type NAME[/FLAVOR]Compose the environment with the named agent harness. Refer to Swapping the agent.
--allow-unsupported-pairingRun even if the resources server does not declare support for the selected agent.
--model / --model-url / --model-api-keyOverride model name, base URL, and API key.
gym env validate my_eval
gym env validate --sync my_eval
gym env validate --environment workplace_assistant
gym env validate --benchmark gsm8k
# or explicit config path(s)
gym env validate --config resources_servers/example_single_tool_call/configs/example_single_tool_call.yaml
# fail unless everything a run needs is defined, credentials included
gym env validate --resources-server example_single_tool_call --model-type openai_model --strict

gym env packages

Each server has its own isolated virtual environment. List the packages installed in a server’s environment.

OptionDescription
--resources-server NAMEName of the resources server.
--search-dir DIRExtra root directory to search for named components. Repeatable.
--outdatedList only outdated packages.
--jsonOutput the package list as JSON.
gym env packages --resources-server example_single_tool_call
# Check for outdated packages
gym env packages \
--resources-server example_single_tool_call \
--outdated

gym env test

Exercise a manifest-backed workload’s verifier fixture in its resources-server environment, or run the existing pytest suite for one or all resources servers.

OptionDescription
NAMEManifest-backed catalog entry whose verifier fixture should run.
--kind environment|benchmarkResolve an ambiguous catalog name.
--update-expectedAtomically update fixture rewards after every behavioral check passes.
--jsonOutput the verifier report as JSON.
--resources-server NAMEResources server to test.
--allTest every resources server. Slow: builds one venv per server, so it must be requested explicitly.
--search-dir DIRExtra root directory to search for named components. Repeatable.
# Exercise a manifest-backed workload verifier
gym env test my_eval
# Test a single server
gym env test --resources-server example_single_tool_call
# Test all servers. Slow: builds one venv per server, so it needs an explicit opt-in.
gym env test --all

gym env publish

Run the local publication check for a manifest-backed workload. The command runs static validation and the verifier fixture, rejects manifest metadata placeholders, and confirms that the exact manifest is discoverable. It includes an experimental annotation only when the manifest’s flag is true. It does not certify the workload, assess implementation quality, require a full evaluation or training run, commit, or push changes.

OptionDescription
NAMEManifest-backed catalog entry to publish.
--kind environment|benchmarkResolve an ambiguous catalog name.
--jsonOutput the publication result as JSON.
gym env publish my_eval
gym env publish my_benchmark --kind benchmark --json

gym env start

Start the NeMo Gym servers (agents, models, resources) defined by the provided configs. Reads configuration from YAML files and runs each configured server in its own environment.

OptionDescription
--config PATHConfig file to load. Repeatable.
--benchmark NAMELoad the named benchmark config (start its servers).
--resources-server NAMELoad the named resources server config.
--model-type NAMELoad the named model server type config.
--search-dir DIRExtra root directory to search for named components. Repeatable.
--agent-type NAME[/FLAVOR]Compose the environment with the named agent harness. Refer to Swapping the agent.
--allow-unsupported-pairingRun even if the resources server does not declare support for the selected agent.
--model, -mServed model identifier. See Selecting the model server.
--model-urlModel server base URL.
--model-api-keyModel server API key.
gym env start \
--resources-server example_single_tool_call \
--model-type openai_model
# Start a benchmark's servers
gym env start \
--benchmark gpqa \
--model-type vllm_model

gym env status

Show all currently running NeMo Gym servers and their health.

OptionDescription
--jsonOutput the server list as JSON.
gym env status
NeMo Gym Server Status:
[1] ✓ example_single_tool_call (resources_servers/example_single_tool_call)
{
'server_type': 'resources_servers',
'name': 'example_single_tool_call',
'port': 58117,
'pid': 89904,
'uptime_seconds': '0d 0h 0m 41.5s',
}
...
3 servers found (3 healthy, 0 unhealthy)

Evaluation

Commands for running evaluations end to end: prepare data, collect rollouts, aggregate sharded runs, verify rollout health, and profile results.

gym eval prepare

Prepare a benchmark’s data by running its prepare.py script and dump the result to disk.

OptionDescription
--config PATHConfig file to load. Repeatable.
--benchmark NAMELoad the named benchmark config.
--search-dir DIRExtra root directory to search for named components. Repeatable.
gym eval prepare --benchmark aime24

gym eval run

Collate data, start the servers, and collect rollouts. This is the main evaluation command. By default it spins up all required servers. Pass --no-serve to collect against servers you already started with gym env start.

OptionDescription
--config PATHConfig file to load. Repeatable.
--benchmark NAMELoad the named benchmark config.
--resources-server NAMELoad the named resources server config.
--model-type NAMELoad the named model server type config.
--search-dir DIRExtra root directory to search for named components. Repeatable.
--no-serveCollect against already-running servers instead of starting them.
--resumeResume from cached rollouts instead of recollecting. Maps to legacy +resume_from_cache=true. Refer to Resume interrupted runs.
--agent-type NAME[/FLAVOR]Compose the environment with the named agent harness. Rejected with --no-serve. Refer to Swapping the agent.
--agent, -aAgent instance to collect rollouts with.
--allow-unsupported-pairingRun even if the resources server does not declare support for the selected agent.
--input, -iInput tasks JSONL file.
--output, -oOutput rollouts JSONL file.
--limitMaximum number of tasks to run.
--num-repeatsRollouts per task (for mean@k metrics). Pass an int to apply to every task, or a dict keyed by agent_ref.name for per-agent counts (e.g. '{simple_agent: 32, swe_agent: 1}') when one input file mixes agents. In dict form, the special key _default is the fallback for agents not explicitly listed; without it, any unlisted row’s agent raises a single consolidated error.
--prompt-configPrompt template YAML to apply.
--concurrencyMaximum number of concurrent samples.
--splitDataset split to use (train, validation, or benchmark). A dataset of the matching type must be declared in the loaded configs. example datasets are smoke-test samples, not a runnable split — run them with --no-serve and --input as shown in the Quickstart.
--model, -mServed model identifier.
--model-urlModel server base URL.
--model-api-keyModel server API key.
--temperatureSampling temperature.
--top-pNucleus sampling top-p.
--max-output-tokensMaximum output tokens.
--disable-aggregationSkip aggregate metrics and the automatic health check. Use for shards that will be combined with gym eval aggregate.
--no-health-checkSkip the automatic post-run health check.
--health-check-workersNumber of rollout-health worker processes. Defaults to the smaller of the CPU count and 8.
--health-check-ignoreComma-separated health-check IDs to exclude from execution and verdict derivation.
# End-to-end: spin up servers, then collect rollouts for a benchmark
gym eval run --benchmark aime24 \
--model-type openai_model \
--output results/aime24.jsonl \
--split validation \
--concurrency 10
# Against an already-running server, with a remote vLLM endpoint
gym eval run --no-serve \
--model-type openai_model \
--resources-server math_with_judge \
--output results/test_001.jsonl \
--split validation \
--model openai/gpt-oss-120b \
--model-url http://0.0.0.0:10240/v1 \
--model-api-key dummy_key \
--temperature 1.0 \
--top-p 1.0
# Per-agent repeats: one input file pins different agents per row via agent_ref.name
gym eval run --no-serve \
--model-type openai_model \
--input mixed_agents.jsonl \
--output results/mixed_rollouts.jsonl \
--num-repeats '{agent_alpha: 4, agent_beta: 1, _default: 1}'

Generation parameters

The most common sampling parameters have dedicated flags on gym eval run: --temperature, --top-p, and --max-output-tokens. These map onto responses_create_params.temperature, responses_create_params.top_p, and responses_create_params.max_output_tokens.

Any other responses_create_params field that has no dedicated flag can be set with a raw Hydra override using the ++responses_create_params.<field> syntax. Overrides are merged into each input row’s existing responses_create_params with a shallow merge (top-level keys only):

gym eval run --no-serve \
--agent example_single_tool_call_simple_agent \
--input weather_query.jsonl \
--output weather_rollouts.jsonl \
--temperature 1.0 \
--top-p 1.0 \
--max-output-tokens 4096 \
++responses_create_params.reasoning.effort=low

Because the merge is shallow, setting a field inside a nested object, such as ++responses_create_params.reasoning.effort=low, replaces the row’s entire nested dictionary at that key. Other fields under the same nested object are not preserved.

Resume interrupted runs

Pass --resume to restart the same command after a crash or interruption and pick up only the rows that have not finished yet.

How it works:

  • Materialized inputs. On the first run, the fully expanded input rows (after --num-repeats, --limit, --prompt-config, and any overrides) are written to a sidecar file next to your output. The path is derived from --output by appending _materialized_inputs to the stem — so rollouts.jsonl produces rollouts_materialized_inputs.jsonl.
  • Incremental output. Successful rollouts are flushed to the main output JSONL after each completion; retriable failures go to a <stem>_failures.jsonl sidecar, so partial progress survives a crash.
  • Failed agent calls. By default a failed agent /run ends the run, because the alternative is a score reported over fewer rollouts than were asked for. Pass +route_failures_to_sidecar=true to let the run continue instead: the call is then recorded in the same sidecar as one attempt, with no reward. Every rollout that leaves the score is logged as it happens, whichever layer classified it, and the totals appear with the closing artifact paths. The class says whether the rollout ran: agent_run_error when the agent itself answered, since a NeMo Gym agent returns 500 when its handler raises, and agent_request_failed for a gateway status (429, 502, 503, 504) or no usable reply, neither of which says anything about the rollout. Other rollouts keep running, the score is computed over completed rollouts only, and resume re-dispatches the row up to NEMO_GYM_MAX_ROLLOUT_ATTEMPTS. A run in which nothing produced a result fails instead of reporting an empty score.
  • Counting failures in the score. Metrics are computed over completed rollouts only. +count_failure_classes_as_zero=[<class>] includes the named sidecar rows in /aggregate_metrics so they land in the denominator: [agent_run_error] counts the rollouts that reached the agent and broke, which is what a model emitting unparsable tool calls looks like from here, and needs +route_failures_to_sidecar=true to have anything to count. A row that carries no reward is scored zero for the metric input only, so neither the sidecar nor the rollouts JSONL records a verdict that no verifier gave. It works the same in gym eval run and gym eval aggregate.
  • Matching. On resume, completed work is matched by (task_index, rollout_index) against the materialized inputs, and already-completed rows are skipped. The run prints a summary such as the number of original input rows, rows already done, and rows that still need to be run.
  • Fallback. If either the materialized inputs or the output file is missing, resume is skipped and the run starts fresh. Without --resume, existing output is cleared before the run.

If you change the config, schema, or data between runs, the materialized inputs become stale and resume will diff against the old expansion. Delete the *_materialized_inputs.jsonl file (and the output file) to start fresh.

gym eval aggregate

Merge sharded rollout results into a single rollouts file with aggregate metrics. Reads every JSONL file matching --input-glob, recomputes aggregate metrics over the global union of records, and writes a <output stem>_aggregate_metrics.json next to the merged rollouts. Use this to combine shards produced by gym eval run --no-serve --disable-aggregation.

OptionDescription
--config PATHConfig file to load. Repeatable.
--input-glob, -iGlob (or comma-separated globs) matching the rollout shards to aggregate.
--output, -oPath for the merged rollouts and aggregate-metrics file.
--no-health-checkSkip the automatic post-aggregation health check.
--health-check-workersNumber of rollout-health worker processes. Defaults to the smaller of the CPU count and 8.
--health-check-ignoreComma-separated health-check IDs to exclude from execution and verdict derivation.
gym eval aggregate \
--config benchmarks/aime24/config.yaml \
--config responses_api_models/vllm_model/configs/vllm_model.yaml \
--input-glob 'results/rollouts-rs*-chunk*.jsonl' \
--output results/rollouts.jsonl

gym eval export

Export supported ng_trajectory attachments from a Gym rollouts JSONL file as ATIF v1.7. The command writes one ATIF trajectory per rollout plus a manifest that preserves Gym’s task and rollout identity. It is a strict offline conversion: if Gym cannot represent a source trajectory completely in the supported ATIF subset, the export fails instead of silently omitting data. All invalid JSONL rows are reported together, and no destination is published if any row fails. Captured model attempts not selected by a turn, such as failed retries, remain in extra.nemo_gym.surplus_model_calls; they do not become canonical ATIF steps or contribute to final usage totals. The destination directory must not already exist.

This command does not parse ATOF and does not load NeMo Relay. Each ATIF trajectory takes its agent name from the source row’s agent_ref.name; --agent-version records the version of that agent implementation.

OptionDescription
--format atifOutput format. ATIF is the only format supported initially.
--rollouts PATHGym rollouts JSONL containing ng_trajectory version 1.0 attachments.
--output-dir DIRNew directory for the ATIF files and manifest.
--session-id IDStable identifier for the source evaluation run.
--agent-version VERSIONVersion of the agent implementation that produced the rollouts.
gym eval export \
--format atif \
--rollouts results/rollouts.jsonl \
--output-dir results/atif \
--session-id eval-2026-08-25 \
--agent-version 1.2.3

See the trajectory capability matrix for the initial conversion boundary.

gym eval health-check

Verify rollout quality for an existing run directory. The command reads <run-dir>/rollouts.jsonl by default and writes quality_summary.json and rollout_verdicts.jsonl in the run directory.

OptionDescription
RUN_DIRDirectory in which to write the health reports.
--rollouts-file PATHRollout JSONL path. Relative paths resolve under RUN_DIR; absolute paths are used as written. Defaults to rollouts.jsonl.
--workersNumber of worker processes. Defaults to the smaller of the CPU count and 8.
--ignore-checks, --ignoreComma-separated health-check IDs to exclude from execution and verdict derivation.
gym eval health-check results/my-run
# Select a nonstandard rollout filename and exclude one known check
gym eval health-check results/my-run \
--rollouts-file evaluator_rollouts.jsonl \
--ignore-checks model_call_missing_token_counts

Refer to Rollout Health Checks for verdict semantics, evidence requirements, the complete check catalog, and report schemas.

gym eval profile

Compute a reward profile from collected rollouts. Outputs per-task statistics such as average reward, standard deviation, min/max, and pass rate, useful for filtering tasks before training by difficulty or variance. Requires rollouts collected with --num-repeats greater than 1.

OptionDescription
--inputsMaterialized inputs JSONL fed to rollout collection.
--rolloutsRollouts JSONL produced by collection.
gym eval profile \
--inputs materialized_inputs.jsonl \
--rollouts rollouts.jsonl

Writes three files next to the rollouts, named from its stem:

FileContents
<rollouts>_reward_profiling.jsonlOne row per task.
<rollouts>_agent_metrics.jsonOne entry per agent.
<rollouts>_repeat_level_metrics.jsonOne entry per repeat.

See Repeat-Level Metrics for the full field list and how to read the variability statistics.

gym eval compare

Compare a baseline eval run against a candidate run and write a report.

The report contains a metric table (the change, plus each side’s value and confidence interval), split into key metrics and all other metrics, and a per-task sample-flips table. Confidence intervals are read from what each run recorded; this command does not yet perform statistical tests or issue a pass/fail verdict.

OptionDescription
--baseline PATHBaseline run’s rollouts JSONL. Its *_aggregate_metrics.json sibling is what gets read.
--candidates PATH[,PATH...]Candidate run’s rollouts JSONL. Comma-separated; one candidate is supported today.
--baseline-agg-metrics PATHBaseline’s aggregate-metrics JSON, when it is not the sibling of --baseline.
--candidates-agg-metrics PATH[,PATH...]Candidates’ aggregate-metrics JSON, in --candidates order, when not siblings of --candidates.
--agent NAMEAgent to compare on both sides. Defaults to every agent present in both runs.
--baseline-agent NAMEAgent to read from the baseline. Takes precedence over --agent.
--candidate-agents NAME[,NAME...]Agent to read from each candidate, in --candidates order. Takes precedence over --agent.
--output-dir DIR, -oWhere to write the report. Defaults to the directory containing the candidate rollouts JSONL.
--report-formatmd, json, or both (default).
gym eval compare \
--baseline runs/baseline/rollouts.jsonl \
--candidates runs/candidate/rollouts.jsonl

Writes into the output directory:

FileContents
compare_report.mdHuman-readable report.
compare_report.jsonMachine-readable result (schema_version "1").

If a run was collected with --disable-aggregation it has no *_aggregate_metrics.json; run gym eval aggregate first, or point at the file directly with --baseline-agg-metrics / --candidates-agg-metrics.

gym eval reverify

Recompute rewards by replaying a stored response through a resources server’s /verify endpoint — without re-running model inference. The command accepts native Gym rollouts or the initial completed, text-only Relay ATIF v1.7 profile. It starts the resources server automatically from the provided config.

Before starting, the command checks each resources server’s GET /reverify_mode response. Resources servers report one of three modes:

ModeMeaning
statelessReverification is safe. The verifier is a pure function of (request body, server config).
unsupportedReverification is not safe. The verifier reads per-rollout session state that is no longer present.
unknownThe server has not declared its reverification behaviour. This is the default. Treated as potentially unsafe.

Servers that report unsupported or unknown cause the command to abort unless --force is passed.

OptionDescription
--config PATHConfig file for the resources server. Repeatable.
--benchmark NAMELoad the named benchmark config.
--resources-server NAMELoad the named resources server config.
--search-dir DIRExtra root directory to search for named components. Repeatable.
--input-format gym|atifSelect native Gym rollout input or the bounded Relay ATIF v1.7 profile. Defaults to gym.
--inputs PATH*_materialized_inputs.jsonl from gym eval run.
--rollouts PATHNative Gym rollouts.jsonl. Required when --input-format gym.
--atif-manifest PATHJSONL manifest mapping each Relay ATIF trajectory to _ng_task_index and _ng_rollout_index, with an optional source SHA-256. Missing hashes produce a warning. Required when --input-format atif.
--output PATH, -oOutput JSONL for recomputed rollouts.
--forceOverride the UNSUPPORTED or UNKNOWN reverify mode guard; output is prefixed with unsafe_.
--overwriteDelete an existing output file instead of raising an error.
--resumeResume a partial run: skip rows already in the output and re-verify only the rest.
--judge-failed-onlyRecover only the rollouts whose judge call failed in the original run (from <rollouts>_failures.jsonl), reusing the stored responses. Skips the reverify-mode guard.
--appendWith --judge-failed-only: append recovered rows to an existing --output instead of a fresh file. Mutually exclusive with --overwrite.
--disable-aggregationSkip aggregate-metrics computation (use when combining shards with gym eval aggregate afterward).

ATIF input requires a stateless resources server and rejects --force, --resume, --judge-failed-only, and --append. It fails closed on unsupported or lossy canonical trajectory shapes, including tool-bearing trajectories routed to a resources server with expose_tools_over_mcp: true, and preserves source identity under _ng_atif_provenance without sending that metadata to the verifier. Provider-native payloads and producer-private status records in optional ATIF extra metadata are not interpreted; their producer owns conversion into canonical ATIF fields. A terminal tool step is supported when every call has a correlated observation result.

# Recompute rewards with an updated verifier config (no model server needed)
gym eval reverify \
--config my_resources_server.yaml \
"++head_server.host=127.0.0.1" \
"++head_server.port=9500" \
--inputs results/rollouts_materialized_inputs.jsonl \
--rollouts results/rollouts.jsonl \
--output results/rollouts_reverified.jsonl
# Override a verifier hyperparameter inline
gym eval reverify \
--config my_resources_server.yaml \
"++mcqa.resources_servers.mcqa.grading_mode=lenient_boxed" \
--inputs results/rollouts_materialized_inputs.jsonl \
--rollouts results/rollouts.jsonl \
--output results/rollouts_lenient.jsonl
# Score a Relay ATIF v1.7 trajectory against a materialized Gym task
gym eval reverify \
--config my_resources_server.yaml \
--input-format atif \
--inputs results/materialized_inputs.jsonl \
--atif-manifest relay-atif/manifest.jsonl \
--output results/relay_atif_reverified.jsonl
# Re-run the judge on only the rollouts whose judge failed (reuses stored responses; no inference)
gym eval reverify --judge-failed-only \
--config my_resources_server.yaml \
--inputs results/rollouts_materialized_inputs.jsonl \
--rollouts results/rollouts.jsonl \
--output results/rollouts_recovered.jsonl
# Force through an UNSUPPORTED or UNKNOWN server (output prefixed with unsafe_)
gym eval reverify \
--config my_stateful_server.yaml \
--inputs results/rollouts_materialized_inputs.jsonl \
--rollouts results/rollouts.jsonl \
--output results/rollouts_reverified.jsonl \
--force

See Reverify Rollouts for a full walkthrough.


Contributor Helpers

gym dev test

Run NeMo Gym’s core unit tests with coverage reporting.

gym dev test

Migrating from the legacy commands

The legacy ng_* and nemo_gym_* are deprecated and will be removed in the future release. Use the tables below to find their gym replacement and update your scripts and workflows.

Command mapping

Legacy commandNew command
ng_helpgym --help
ng_versiongym --version
ng_list_benchmarksgym list benchmarks
ng_rungym env start
ng_statusgym env status
ng_dump_configgym env resolve
ng_pip_listgym env packages
ng_init_resources_servergym env init
ng_testgym env test
ng_test_allgym env test --all
ng_prepare_benchmarkgym eval prepare
ng_e2e_collect_rolloutsgym eval run
ng_collect_rolloutsgym eval run --no-serve
ng_aggregate_rolloutsgym eval aggregate
ng_reward_profilegym eval profile
ng_prepare_datagym dataset collate
ng_materialize_promptsgym dataset render
ng_upload_dataset_to_hfgym dataset upload
ng_upload_dataset_to_gitlabgym dataset upload --storage gitlab
ng_download_dataset_from_hfgym dataset download
ng_download_dataset_from_gitlabgym dataset download --storage gitlab
ng_gitlab_to_hf_datasetgym dataset migrate
ng_delete_dataset_from_gitlabgym dataset rm
ng_dev_testgym dev test
ng_reinstalluv sync --extra dev

Replacing Hydra overrides with flags

The most common Hydra overrides now have dedicated flags:

Legacy Hydra overrideNew flag
"+config_paths=[a.yaml,b.yaml]"--config a.yaml --config b.yaml
+agent_name=...--agent
+input_jsonl_fpath=...--input
+input_glob=...--input-glob
+output_jsonl_fpath=...--output
+limit=...--limit
+num_repeats=...--num-repeats
+num_samples_in_parallel=...--concurrency
++split=...--split
++policy_model_name=...--model
++policy_base_url=...--model-url
++policy_api_key=...--model-api-key
++responses_create_params.temperature=...--temperature
++responses_create_params.top_p=...--top-p
++responses_create_params.max_output_tokens=...--max-output-tokens
+mode=...--mode
+output_dirpath=...--output-dir
+should_download=true--download
+prompt_config=...--prompt-config
+resume_from_cache=true--resume
+dataset_name=...--name
+repo_id=...--repo-id
+revision=... / +version=...--revision
+artifact_fpath=...--artifact
+output_fpath=...--output
+create_pr=true--create-pr
+outdated=true--outdated

Common workflows, before and after

# Start servers
# Before:
ng_run "+config_paths=[resources_servers/example_single_tool_call/configs/example_single_tool_call.yaml,responses_api_models/openai_model/configs/openai_model.yaml]"
# After:
gym env start \
--resources-server example_single_tool_call \
--model-type openai_model
# End-to-end rollout collection
# Before:
config_paths="responses_api_models/openai_model/configs/openai_model.yaml,resources_servers/math_with_judge/configs/math_with_judge.yaml"
ng_e2e_collect_rollouts "+config_paths=[${config_paths}]" \
++output_jsonl_fpath=results/aime24.jsonl \
++split=validation
# After:
gym eval run \
--model-type openai_model \
--resources-server math_with_judge \
--output results/aime24.jsonl \
--split validation
# Collect against already-running servers
# Before:
ng_collect_rollouts +agent_name=example_single_tool_call_simple_agent \
+input_jsonl_fpath=weather_query.jsonl \
+output_jsonl_fpath=weather_rollouts.jsonl \
+num_repeats=4 +num_samples_in_parallel=10
# After:
gym eval run --no-serve \
--agent example_single_tool_call_simple_agent \
--input weather_query.jsonl \
--output weather_rollouts.jsonl \
--num-repeats 4 \
--concurrency 10
# Reward profiling
# Before:
ng_reward_profile +input_jsonl_fpath=materialized_inputs.jsonl +rollouts_jsonl_fpath=rollouts.jsonl
# After:
gym eval profile \
--inputs materialized_inputs.jsonl \
--rollouts rollouts.jsonl

Run any command with --help to see its full set of flags, and gym --help to list every group.