CLI Commands
This page documents the NeMo Gym command-line interface.
All functionality is exposed through a single gym entry point, organized into command groups (gym <group> <command>). ng is a drop-in alias for gym. Every group and command supports -h/--help.
The legacy ng_* / nemo_gym_* commands (such as ng_run or nemo_gym_collect_rollouts) still work but are deprecated. Each one prints a notice pointing at its gym replacement and then runs it. See Migrating from the legacy commands at the bottom of this page.
Quick Reference
Common Options
These options are shared across many commands.
The --benchmark, --resources-server, --model-type, and --agent-type selectors resolve a component name to its config file for you, so you do not need to know the project’s directory layout. If you mistype a name, the CLI suggests the closest match. To point at a config file directly, use --config <path> instead.
Selecting the model server
Commands that need a model (gym env start, gym eval run) configure it through four flags:
Swapping the agent
--agent-type NAME runs an environment or benchmark with a different agent harness, without editing any
config. It resolves to responses_api_agents/NAME/configs/NAME.yaml; use NAME/FLAVOR to pick a
non-default config in that directory.
Available on gym env start, gym env validate, gym env prefetch, and gym eval run. Like the other
selectors it is pure shorthand, so --agent-type terminus_2_sandboxed_agent and --config <that file>
behave identically.
Run gym list agents to see the harnesses you can pass.
The composed server instance is renamed after the agent that runs it, so metrics and gym env status
report the harness actually used: composing terminus_2_sandboxed_agent onto
terminal_bench_2_1_opencode_sandboxed_agent produces terminal_bench_2_1_terminus_2_sandboxed_agent.
gym env start prints the instance list, which matters for the two-step flow:
--agent-type chooses the harness to compose; --agent names the running instance to collect rollouts
from. Because the harness is settled once the servers are up, --agent-type is rejected with
--no-serve.
Anything that routes rows — --agent, +agent_map values, +fan_out entries — must name the composed
instance, and Gym says so rather than failing later:
Rows collected before the swap are re-routed for you, so existing rollout files stay usable.
Some tasks only work with the harness they were built for. A resources server may pin the harnesses
required for the task with allowed_agents, and
Gym refuses a pairing it does not list, before any server starts:
Results from an undeclared pairing might not be valid.
Hydra overrides (escape hatch)
The gym CLI is a thin wrapper over Gym’s Hydra config system. Standard flags (--flag value) cover the common inputs. Anything not covered by a flag can still be passed as a raw Hydra override using +key=value (add a new key) or ++key=value (add or override an existing key). Unknown overrides are forwarded to Hydra untouched, so advanced config composition is fully functional.
When a command needs overrides for keys that have no dedicated flag, you can keep the whole command in Hydra form (config paths included) or mix --flag and +key=value styles in one invocation:
General
gym --help
List all command groups and their descriptions.
gym --version
Print the NeMo Gym version along with Python, key dependency, and system information.
Discovery
Commands for discovering benchmarks, environments, model servers, agent harnesses, and resources servers.
gym list <type> lists all components of that type; gym list <type> <name> inspects that one.
All commands accept:
--jsonwhich prints machine-readable output instead of a human-readable table.--search-dirwhich allows you to surface your own Gym-compatible components alongside the built-ins.
gym list benchmarks
List the benchmarks available in NeMo Gym, with their kind, manifest status, domain, and description.
The value in the Name column is the benchmark identifier and can be passed to --benchmark.
To see which agent harness a benchmark composes with, inspect it with gym list benchmarks <name>.
gym list environments
List the manifest-first catalog unioned with legacy environment and benchmark configs. Manifest-backed entries expose their declared metadata; unmigrated entries remain visible with no-manifest status.
gym list agents
List the agent harnesses under responses_api_agents/, with each one’s composition pattern — composable (Pattern A: references a separate resources server, so it can be wired into a matching environment) vs self-contained (Pattern B: ships its own framework/environment) — plus its config variants.
gym list models
List the model servers under responses_api_models/.
The value in the Model column corresponds to a config flavor and can be passed as the --model-type flag to other commands.
gym list resources-servers
List the resources servers under resources_servers/.
The value in the Name column corresponds to the config flavor and can be passed as --resources-server.
gym search
Filter a listing to fuzzy matches on a query (against each component’s name/value, description, domain, and type-specific fields). Without a type, the command searches the unified environment and benchmark catalog.
Datasets
Commands for preparing, previewing, and managing datasets. By default dataset transfer commands use HuggingFace; pass --storage gitlab to target the GitLab Registry.
gym dataset upload
Upload a prepared local JSONL dataset to HuggingFace (default) or GitLab.
gym dataset download
Download a dataset from HuggingFace (default) or GitLab.
gym dataset rm
Delete a dataset from the GitLab Registry. Prompts for confirmation.
gym dataset migrate
Migrate a JSONL dataset to HuggingFace from GitLab. Use gym dataset upload if you do not want automatic GitLab deletion.
gym dataset collate
Validate and collate a dataset, generating metrics and statistics.
gym dataset render
Generate a dataset preview by materializing prompts from a raw input file and a prompt template, producing JSONL with populated responses_create_params.input for RL training.
Each input row must not already have a populated responses_create_params.input; the command applies the prompt template from --prompt-config to each row, fills in the input, and preserves the row’s other fields.
Which data-preparation command should I use?
gym dataset render— a focused, standalone step that applies a prompt template to raw rows to populateresponses_create_params.input. No servers are started. Use it when you have raw data and just need to turn it into prompt-ready rows.gym dataset collate— the full preparation pipeline for training: it can download missing datasets, validate data, and compute dataset metrics, writing train/validation splits and metrics artifacts. Use it to prepare and validate datasets for training or PR submission.
Environments
Commands for developing, running, and inspecting environments. An environment defines tasks, verification, and its interaction surface; Gym config composes it with agent, model, and runtime components.
gym env init
Scaffold a resources server, or create a manifest-backed environment or benchmark skeleton. Workload scaffolds include a manifest, Gym config, sample data, and the extension points selected by the integration profile.
gym env resolve
Resolve the configs, flags, and overrides into a final merged config and print it. Useful for debugging configuration. Secrets are hidden.
Unlike gym env validate, resolve substitutes no dummy model, so a model config’s policy_* values must be supplied for its interpolations to resolve.
gym env validate
Validate a manifest-backed workload or legacy config without starting Ray or server subprocesses. Manifest validation resolves the Gym config offline, checks composition mirrors and datasets, and reports the resolved components. Legacy validation retains the existing config pre-flight behavior. Exits 0 when valid, or 1 with a clean message when not.
By default, values a run needs but the config leaves undefined (such as policy_base_url from env.yaml) are filled with placeholders so the rest can still be checked, and reported as a warning. Pass --strict to fail on them instead.
gym env packages
Each server has its own isolated virtual environment. List the packages installed in a server’s environment.
gym env test
Exercise a manifest-backed workload’s verifier fixture in its resources-server environment, or run the existing pytest suite for one or all resources servers.
gym env publish
Run the local publication check for a manifest-backed workload. The command runs static validation and the verifier fixture, rejects manifest metadata placeholders, and confirms that the exact manifest is discoverable. It includes an experimental annotation only when the manifest’s flag is true. It does not certify the workload, assess implementation quality, require a full evaluation or training run, commit, or push changes.
gym env start
Start the NeMo Gym servers (agents, models, resources) defined by the provided configs. Reads configuration from YAML files and runs each configured server in its own environment.
gym env status
Show all currently running NeMo Gym servers and their health.
Evaluation
Commands for running evaluations end to end: prepare data, collect rollouts, aggregate sharded runs, verify rollout health, and profile results.
gym eval prepare
Prepare a benchmark’s data by running its prepare.py script and dump the result to disk.
gym eval run
Collate data, start the servers, and collect rollouts. This is the main evaluation command. By default it spins up all required servers. Pass --no-serve to collect against servers you already started with gym env start.
Generation parameters
The most common sampling parameters have dedicated flags on gym eval run: --temperature, --top-p, and --max-output-tokens. These map onto responses_create_params.temperature, responses_create_params.top_p, and responses_create_params.max_output_tokens.
Any other responses_create_params field that has no dedicated flag can be set with a raw Hydra override using the ++responses_create_params.<field> syntax. Overrides are merged into each input row’s existing responses_create_params with a shallow merge (top-level keys only):
Because the merge is shallow, setting a field inside a nested object, such as ++responses_create_params.reasoning.effort=low, replaces the row’s entire nested dictionary at that key. Other fields under the same nested object are not preserved.
Resume interrupted runs
Pass --resume to restart the same command after a crash or interruption and pick up only the rows that have not finished yet.
How it works:
- Materialized inputs. On the first run, the fully expanded input rows (after
--num-repeats,--limit,--prompt-config, and any overrides) are written to a sidecar file next to your output. The path is derived from--outputby appending_materialized_inputsto the stem — sorollouts.jsonlproducesrollouts_materialized_inputs.jsonl. - Incremental output. Successful rollouts are flushed to the main output JSONL after each completion; retriable failures go to a
<stem>_failures.jsonlsidecar, so partial progress survives a crash. - Failed agent calls. By default a failed agent
/runends the run, because the alternative is a score reported over fewer rollouts than were asked for. Pass+route_failures_to_sidecar=trueto let the run continue instead: the call is then recorded in the same sidecar as one attempt, with no reward. Every rollout that leaves the score is logged as it happens, whichever layer classified it, and the totals appear with the closing artifact paths. The class says whether the rollout ran:agent_run_errorwhen the agent itself answered, since a NeMo Gym agent returns 500 when its handler raises, andagent_request_failedfor a gateway status (429, 502, 503, 504) or no usable reply, neither of which says anything about the rollout. Other rollouts keep running, the score is computed over completed rollouts only, and resume re-dispatches the row up toNEMO_GYM_MAX_ROLLOUT_ATTEMPTS. A run in which nothing produced a result fails instead of reporting an empty score. - Counting failures in the score. Metrics are computed over completed rollouts only.
+count_failure_classes_as_zero=[<class>]includes the named sidecar rows in/aggregate_metricsso they land in the denominator:[agent_run_error]counts the rollouts that reached the agent and broke, which is what a model emitting unparsable tool calls looks like from here, and needs+route_failures_to_sidecar=trueto have anything to count. A row that carries no reward is scored zero for the metric input only, so neither the sidecar nor the rollouts JSONL records a verdict that no verifier gave. It works the same ingym eval runandgym eval aggregate. - Matching. On resume, completed work is matched by
(task_index, rollout_index)against the materialized inputs, and already-completed rows are skipped. The run prints a summary such as the number of original input rows, rows already done, and rows that still need to be run. - Fallback. If either the materialized inputs or the output file is missing, resume is skipped and the run starts fresh. Without
--resume, existing output is cleared before the run.
If you change the config, schema, or data between runs, the materialized inputs become stale and resume will diff against the old expansion. Delete the *_materialized_inputs.jsonl file (and the output file) to start fresh.
gym eval aggregate
Merge sharded rollout results into a single rollouts file with aggregate metrics. Reads every JSONL file matching --input-glob, recomputes aggregate metrics over the global union of records, and writes a <output stem>_aggregate_metrics.json next to the merged rollouts. Use this to combine shards produced by gym eval run --no-serve --disable-aggregation.
gym eval export
Export supported ng_trajectory attachments from a Gym rollouts JSONL file as ATIF v1.7. The command writes one ATIF
trajectory per rollout plus a manifest that preserves Gym’s task and rollout identity. It is a strict offline conversion:
if Gym cannot represent a source trajectory completely in the supported ATIF subset, the export fails instead of silently
omitting data. All invalid JSONL rows are reported together, and no destination is published if any row fails. Captured
model attempts not selected by a turn, such as failed retries, remain in extra.nemo_gym.surplus_model_calls; they do not
become canonical ATIF steps or contribute to final usage totals. The destination directory must not already exist.
This command does not parse ATOF and does not load NeMo Relay. Each ATIF trajectory takes its agent name from the source
row’s agent_ref.name; --agent-version records the version of that agent implementation.
See the trajectory capability matrix for the initial conversion boundary.
gym eval health-check
Verify rollout quality for an existing run directory. The command reads <run-dir>/rollouts.jsonl by default and writes
quality_summary.json and rollout_verdicts.jsonl in the run directory.
Refer to Rollout Health Checks for verdict semantics, evidence requirements, the complete check catalog, and report schemas.
gym eval profile
Compute a reward profile from collected rollouts. Outputs per-task statistics such as average reward, standard deviation, min/max, and pass rate, useful for filtering tasks before training by difficulty or variance. Requires rollouts collected with --num-repeats greater than 1.
Writes three files next to the rollouts, named from its stem:
See Repeat-Level Metrics for the full field list and how to read the variability statistics.
gym eval compare
Compare a baseline eval run against a candidate run and write a report.
The report contains a metric table (the change, plus each side’s value and confidence interval), split into key metrics and all other metrics, and a per-task sample-flips table. Confidence intervals are read from what each run recorded; this command does not yet perform statistical tests or issue a pass/fail verdict.
Writes into the output directory:
If a run was collected with --disable-aggregation it has no *_aggregate_metrics.json; run gym eval aggregate first, or point at the file directly with --baseline-agg-metrics / --candidates-agg-metrics.
gym eval reverify
Recompute rewards by replaying a stored response through a resources server’s /verify endpoint — without re-running model inference. The command accepts native Gym rollouts or the initial completed, text-only Relay ATIF v1.7 profile. It starts the resources server automatically from the provided config.
Before starting, the command checks each resources server’s GET /reverify_mode response. Resources servers report one of three modes:
Servers that report unsupported or unknown cause the command to abort unless --force is passed.
ATIF input requires a stateless resources server and rejects --force, --resume, --judge-failed-only, and --append. It fails closed on unsupported or lossy canonical trajectory shapes, including tool-bearing trajectories routed to a resources server with expose_tools_over_mcp: true, and preserves source identity under _ng_atif_provenance without sending that metadata to the verifier. Provider-native payloads and producer-private status records in optional ATIF extra metadata are not interpreted; their producer owns conversion into canonical ATIF fields. A terminal tool step is supported when every call has a correlated observation result.
See Reverify Rollouts for a full walkthrough.
Contributor Helpers
gym dev test
Run NeMo Gym’s core unit tests with coverage reporting.
Migrating from the legacy commands
The legacy ng_* and nemo_gym_* are deprecated and will be removed in the future release.
Use the tables below to find their gym replacement and update your scripts and workflows.
Command mapping
Replacing Hydra overrides with flags
The most common Hydra overrides now have dedicated flags:
Common workflows, before and after
Run any command with --help to see its full set of flags, and gym --help to list every group.