Release Notes
v0.6.0
Release Summary
NeMo Gym v0.6.0 lets you use external agent harnesses during RL training and adds Switchyard for evaluating model-routing strategies.
Highlights:
- Use supported external agent harnesses during RL training; NeMo Gym preserves each model call’s token IDs and reconstructs the multi-step run as one training response
- Run routing-aware evaluations with Switchyard by comparing fixed and routed conditions on the same benchmark
- Trust and debug rollouts with automatic health checks, end-to-end traces, and per-agent token, turn, tool-call, and latency diagnostics
- Evaluate and compare multiple agents and datasets in one run, with task-level routing to the appropriate harness
- Submit eval jobs to a Slurm cluster directly from NeMo Gym, scaling vLLM across GPUs or nodes for higher rollout concurrency
First-Time Contributors
We welcomed 19 new contributors to NeMo Gym with this release:
- @giuliolovisotto added automatic rollout health checks and documented how to diagnose rollout quality
- @mlazuka added per-rollout and per-run aggregate observability metrics
- @imxj added the LiteLLM model server
- @max-sudolabs added the E2B sandbox provider
- @aroshanghias-nvd added a simple agent with configurable context-compaction policies
- @ehosseiniasl added video input support to the vLLM model server
- @adv-andrew added the synthetic OpenAir 5G congestion-control environment
- @knayaka added the Citation IF instruction-following benchmark
- @nickson-quak added a social-bias evaluation environment that separately scores answer correctness and whether explanations rely on stereotypes
- @gnalbandyan added the multimodal HLE vision benchmark
- @michal2409 improved OpenHands reliability and generation-limit handling for SWE agents
Thank you to all 64 NeMo Gym contributors this cycle, including 19 first-time contributors!
Configure Tasks and Data
- Resources servers can declare their datasets, and
gym dataset collaterecords each row’stask_sourceso NeMo Gym can select the correct agent at run time - Use
agent_mapto reroute selected tasks or fan-out to run the same tasks with multiple agents; invalid agent mappings fail before execution - Each resources server defines a typed
task_dataschema. Collation validates rows before a run, andgym env schema --resources-server <name>shows the expected data shape
Configure Agent Harnesses
- Use
--agent-typeto swap a benchmark or environment’s configured harness, with compatibility checks for unsupported pairings - OpenCode Sandboxed Agent runs OpenCode inside task-specific NeMo Gym sandboxes, with configurations for SWE-Bench Verified and Multilingual, DeepSWE, and Terminal-Bench 2.1
- Cline Agent runs the Cline CLI headlessly and converts its tool-use trace into a NeMo Gym trajectory; this initial integration is evaluation-only
- Conversational Tool Use workflow adds three agents that generate customer-service domains, policies, tools, and scenarios, plus a policy-loop agent and resources server that simulate and score multi-turn customer and tool interactions
- OSWorld Agent runs OSWorld desktop tasks end to end, executes model-generated GUI actions, and returns NeMo Gym trajectories with OSWorld rewards
- Prime Agent runs Prime Intellect’s agent headlessly through NeMo Gym, with included math and Reasoning Gym environments
- Simple Strands Agent wraps the Simple Strands Resolver agent with configurable local tools; included environments cover math and Reasoning Gym
- Terminus-2 Agent runs Harbor’s terminal-command agent through NeMo Gym, with AnyTerminal integration and support for other terminal-compatible resources servers
- Simple Agent with Compaction applies configurable history policies before each model call, including policies that omit older image observations or reasoning blocks
- Image Tools Agent lets models iteratively transform and inspect images, scores their image-tool calls, and delegates final-answer verification to the mapped environment
- Verified Code QA (VCQA) Agent gives models read-only repository snapshots or Git histories to investigate with file, search, and shell tools, then scores answers against a rubric with an LLM judge
- OpenClaw with AnyTerminal runs the evaluation-only OpenClaw harness inside Terminal Bench task containers and scores each run with the task’s tests
Configure Models
- LiteLLM Model connects NeMo Gym to providers supported by LiteLLM
- Switchyard Model compares model-routing strategies on the same benchmark and records the route and configuration used for reproducible results
- The vLLM model server now supports video inputs
- Use
sampling_overridesto enforce generation settings when external agent harnesses omit them, preserving on-policy training behavior - Use
endpoint_fileto update a vLLM backend address without restarting NeMo Gym, with bounded connection retries during endpoint rotation
Rollout Observability and Export
- Export configs, metrics, and rollouts through a shared exporter framework. MLflow is new; W&B uses the same path
- Automatic rollout health checks flag missing turns, inconsistent model calls, token-count mismatches, and runaway generation, with structured reports from evaluation and aggregation
- NeMo Lens telemetry traces each rollout across NeMo Gym’s agent, model, and resources servers, making latency and failures easier to diagnose
- Rollout diagnostics report token usage, turns, tool calls, and latency, with agent-level traces for OpenClaw, Hermes, Pi, OpenCode, and SWE agents
- Repeated-rollout metrics now include variability and 95% confidence intervals
Evaluation and Training
- Compatible external agent harnesses can produce on-policy training data: NeMo Gym preserves exact token IDs across model calls and rebuilds each multi-step episode as one training response
- Opt in to
route_failures_to_sidecarto keep an evaluation running when agent calls fail; coverage reporting identifies excluded rollouts, and selected failure classes can instead count as zero - Trust eval artifacts: automatic health checks catch missing turns, broken model calls, and runaway generation; rollouts record tokens, turns, tool calls, and latency
- Verifiers can return a human-readable
failure_reasonto explain unsuccessful rollouts
Sandboxing and Orchestration
- E2B joins the built-in sandbox providers with template building, command execution, file operations, and lifecycle management
- The OpenSandbox provider adds asynchronous interactive PTY sessions and detached PTY execution for long-running commands
- OpenSandbox adds configurable network policy, run-scoped cleanup, bounded background-status polling, and an OpenSandbox backend for OSWorld
- The Apptainer provider accepts a custom binary path
- Experimental
gym eval submitruns one or more vLLM instances on a Slurm node or distributes whole replicas across nodes, with configurable environment variables, container mounts, and outputs
Environments and Benchmarks
- Software engineering and agentic workflows: Sandboxed OpenCode configurations for SWE-Bench Verified and Multilingual repository repair, DeepSWE software-engineering tasks, and Terminal-Bench 2.1 terminal tasks; Verified Code QA for repository investigation; OSWorld for desktop automation; and synthetic conversational tool use
- Finance: Vals Finance Agent v2 for multi-step financial analysis and local full-text search over SEC filings
- Long context, knowledge, and reasoning: LongMemEval for conversational memory, RULER v2 for long-context reasoning, SpartQA for spatial reasoning, and new math and Reasoning Gym pairings across agent harnesses
- Safety and instruction following: AgentIF for agentic instruction-following scenarios, Citation IF for citation compliance, and synthetic social-bias questions that separately score answer correctness and whether explanations rely on stereotypes
- Science and multimodal: Expanded Litmus ADME drug-discovery profiles, HLE vision for multimodal expert-level questions, image-tool PivotRL verification, and a synthetic OpenAir 5G congestion-control environment
- Aviary environments: Standalone entries for scientific data analysis with BixBench-Hypothesis and BixBench, grade-school math with GSM8K, and multi-hop question answering with HotpotQA
See the Available Environments table for the full list.
Deprecation Notices
upload_rollouts_to_wandbhas been renamed toupload_rolloutsbecause it now controls rollout uploads to every configured exporter. The old name still works but emits a deprecation warning- NeMo Gym now requires
openai==2.44.0instead ofopenai<=2.7.2. A parent environment that forces another version fails configuration by default. Setallow_openai_version_skew: trueonly when intentionally allowing the parent and server environments to use different SDK versions +agent_name=<name>now sends every row to that agent. Previously, it only applied to rows that did not specify an agent. Use+agent_map=...to reroute only selected rows- Datasets collated before v0.6.0 may contain
agent_refwithouttask_source. They still run, but this routing format is deprecated; rungym dataset collateagain or setagent_mapexplicitly - Because collation now validates task data, malformed or mismatched rows that were previously accepted may fail during preparation. Use
gym env schema --resources-server <name>to see the expected shape - Agentic math environments now use harness-first names. Previous environment names, config paths, preparation scripts, and agent references remain supported as deprecated aliases and emit migration warnings:
reasoning_gym_claude_code→claude_code_reasoning_gymreasoning_gym_hermes→hermes_reasoning_gymreasoning_gym_orchestrator→langgraph_orchestrator_reasoning_gymreasoning_gym_parallel_thinking→langgraph_parallel_thinking_reasoning_gymreasoning_gym_reflection→langgraph_reflection_reasoning_gymreasoning_gym_rewoo→langgraph_rewoo_reasoning_gym
- The following agent config files were renamed. Previous config paths and
--agent-typeflavors remain supported as deprecated aliases and emit migration warnings:stirrup_gdpval.yaml→stirrup_agent.yamltau2_agent.yaml→tau2.yamlacereason-math.yaml→verifiers_agent.yaml
Bug Fixes
- Converting between Responses and Chat Completions now preserves structured tool outputs and reasoning metadata and supports audio, file, custom-tool, and parallel tool-call inputs
- Fixed recursive
config_pathsoverrides so nested configurations apply in the intended order gym eval run --splitnow fails immediately when no declared dataset matches the requested split- Dry runs now fail when server environment setup fails and report unresolved configuration values and catalog lookups clearly
- Improved OpenSandbox reliability around authentication, startup races, unavailable backends, out-of-memory failures, and cleanup
- Added support for running Docker sandboxes as a non-root user
- Preserved exact vLLM token metadata instead of re-tokenizing prompts
- Improved SWE agent reliability across OpenHands, OpenCode, and SWE-Bench
- Fresh runs now clear outdated failure reports when reusing an output path
- Fixed AnyTerminal secret redaction so unrelated configuration values are not modified
- OpenClaw now preserves partial execution traces when a run times out
Documentation
- Added reproducible Nemotron 3.5 Lightning evaluation recipe across agentic, knowledge, reasoning, and science benchmarks
- Added a guide for diagnosing rollout quality with
gym eval health-check - Added a tutorial for using external agent harnesses to generate on-policy training rollouts.
- Added a setup guide for the E2B sandbox provider
- Added a guide for running routing-aware evaluations with Switchyard and comparing results across routing conditions
- Added practical guidance for verifying equivalent answers
v0.5.1
Patch release focused on security, container reliability, and legal/attribution updates — no breaking changes from v0.5.0.
- Security: removed bundled royalty-bearing codec binaries and dependencies flagged in legal review; bumped
orjson,msgpack, andpython-multipartfor CVE remediation - Docker: installed
enrootand declared abashentrypoint in the Gym container; installednvidia-container-toolkitso enroot’s NVIDIA GPU hook works inside the container; documented the pre-built image’sdocker runusage - Legal: expanded
ATTRIBUTIONS.mdto a full 167-package inventory; added the PinchBench modified-component notice and NVIDIA disclaimer - Documentation: fixed stale Python version references (3.12 → 3.13.14) across install and setup guides; cherry-picked the model-call capture guide and tutorial path redirects from v0.5.0
v0.5.0
Release Summary
The NeMo Gym 0.5.0 release expands the sandbox ecosystem to seven providers, adds four new general-purpose agent harnesses (Codex, KiloCode, RemoteAgent, and Any-SWE) bringing the total to 20, adds 21 new benchmarks and environments, and wires rollout observability end-to-end from the model server boundary through agent transcripts.
Highlights:
- Seven sandbox providers: Docker, Daytona, ECS Fargate, Enroot, and OpenShell join OpenSandbox and Apptainer; large-scale OpenSandbox reliability significantly improved
- Four new agent harnesses: Codex CLI, KiloCode, RemoteAgent, and
anyswe_agent - Recompute rewards from stored rollouts without re-running inference with
gym eval reverify - Rollout observability joined end-to-end: model-call capture, agent observations, and a standardized
ng_trajectoryschema - 21 new environments across six domains: Agentic, Knowledge and instruction following, Long context, Science and coding, Translation and multilingual, and Reasoning
First-Time Contributors
We welcomed 22 new contributors to NeMo Gym with this release:
- @mpatel31415 added
gym eval reverifycommand and--judge-failed-onlyflag for recovering failed judge rows - @Glorf added per-rollout model-call capture, Docker and ECS Fargate sandbox providers, rollout observation contract, Claude Code rollout observations, and standardized
ng_trajectoryschema - @nblintao added per-request policy endpoint override for the SWE agents, enabling RL training frameworks to route each episode through a per-episode recording proxy
- @JeffPengCoder brought OSWorld — a stateful desktop GUI benchmark — into the benchmark catalog
- @rystewart-nvidia added the Legal Agent Bench integration, exposing Harvey’s 1,749-task LAB benchmark through standard Gym eval commands
- @fallintoplace fixed the long-standing disagreement between runtime aggregate metrics and the persisted
output.jsonl - @jonathanlli added RULER pretrain evaluation, enabling text-completion scoring for base and midtraining checkpoints
- @thompsonb overhauled WMT24++ and FLORES translation evaluation (55 locales, 219 language pairs, chrF/spBLEU scoring); expanded MMLU-ProX to 29 languages
- @hkumar92 added the PinchBench agentic benchmark
- @pachmu added ToolSandbox, IHEval, RoleMRC, and RAGTruth benchmarks
Thank you to all 57 NeMo Gym contributors this cycle, including 22 first-time contributors!
Command Line Interface
gym eval reverify— re-run only the verifier on stored rollouts;--judge-failed-onlyrecovers rows that failed due to a flaky judge without re-verifying successful rolloutsgym listandgym searchextended to cover models, resources-servers, and agents;gym list <type> <name>drills into a single artifact- External plugins discoverable via
--search-dirand environment variables;-v/--verbosenow accepted before any subcommand
Sandboxing
Five new built-in providers: Docker, Daytona, ECS Fargate, Enroot, and OpenShell join OpenSandbox and Apptainer. Provider choice is a one-line config swap — any agent built on nemo_gym.sandbox works with any provider unchanged.
OpenSandbox reliability at scale is significantly improved: keepalive-bounded transport eliminates silent rollout zeroing at concurrency 300–1500; image registry auth supports private container images; sandbox resources are automatically labeled with team, user, and workload identifiers.
See Available Sandbox Providers for the full list.
Configure Agent Harnesses
New harnesses join the existing set (Claude Code, Hermes, mini-SWE-Agent, OpenClaw, Pi, and more):
- Codex and KiloCode integrate
codex execandkilo runrespectively, routing model calls through Gym for per-rollout capture anyswe_agentruns any Gym harness inside a SWE task containerswe_agentsadds OpenCode as a supported agent framework alongside OpenHands, with DeepSWE and DeNovoSWE dataset support and message replay for trajectory branching- RemoteAgent drives any external service that implements
POST /v1/responses, with Gym owning the tool loop and verification
Configure Models
- vLLM can now drive
/v1/completionsfor base and pretrain checkpoint evaluation via opt-inuse_completions_api - All Gym model servers now accept
stream: trueon/v1/chat/completionsvia synthesized SSE, unblocking streaming-first clients such as OpenClaw and Codex - Add
expose_tools_over_mcp: trueto any resources server config to serve its tools over MCP with no handler code changes
Rollout Observability
- Per-rollout model-call capture records requests, responses, token usage, and latency at the model server boundary
- Claude Code transcripts populate
ng_agent_observations; agent observations and model-call capture are joined through a standardizedng_trajectoryschema - Judge failures are routed to a
_failures.jsonlsidecar, keeping aggregate metrics over successfully-judged rows only
New Benchmarks and Environments
21 new environments across six domains:
- Agentic: PinchBench (147 real-world tasks), OSWorld (desktop GUI with VM-backed evaluation), Legal Agent Bench (1,749 Harvey LAB tasks), ToolSandbox (Apple multi-turn tool-use), BrowseComp (web research), BioMNIBench DA, Tau3 banking (BM25+grep offline eval path)
- Knowledge and instruction following: SECQUE, FinanceBench, Finance SEC Search, IHEval (instruction hierarchy, rule-based), Litmus-Bench v0.1, RoleMRC (role-play MRC), RAGTruth (hallucination detection)
- Long context: NIAH (retrieval with overlap penalty)
- Science and coding: CVDP Agentic (expanded to support the agentic subset, harness-agnostic)
- Translation and multilingual: WMT24++ (expanded from 5 to 55 locales), FLORES (expanded from 30 to 219 language pairs, chrF/spBLEU scoring), MMLU-ProX (expanded to 29 languages); RULER now supports pretrain text-completion evaluation
- Reasoning: ReasoningGym environments — six agentic variants: Claude Code, Hermes, and four LangGraph-based variants (orchestrator, reflection, parallel thinking, and ReWOO)
See the Available Environments table for the full list.
Deprecation Notices
- WMT24++ and FLORES scores from prior versions are not comparable with this version’s chrF/spBLEU output
- Python 3.13.14 is now required (previously 3.12); users running Gym in Python 3.12 environments must upgrade
Bug Fixes
- SciCode realigned to the AA 65-problem test set with per-rollout subtask accuracy reporting
- Fixed silent rollout zeros at high concurrency against OpenSandbox (keepalive-bounded transport)
- Fixed frozen rollouts in long-running benchmarks (TCP keepalive on global aiohttp connector)
- Fixed
gym eval runfailing withFileNotFoundErrorwhen the output directory did not exist - Fixed
tool_choicesent to vLLM withouttools, causing request rejection - Fixed Claude Code
max_turnshardcoded to 30;max_turns: nullnow removes the cap - Fixed Apptainer sandbox env vars injected into subprocess argv instead of environment
- Fixed MCQA answer parsing for wrapped formats (
$D$,(D),\boxed{\text{Answer: G}}) - Fixed aggregate metrics including non-persisted rollouts, causing disagreement with
output.jsonl
Documentation
- Rewrote the key terminology glossary with a Gym overview, component map, and links to how-to pages
- New page documenting the Anthropic Messages dialect (
POST /v1/messages) and wiring Claude Code through a Gym model server - Updated NeMo RL v0.7.0 compatibility guidance
- Migrated all remaining docs examples from legacy
ng_run/ng_collect_rolloutsto the unifiedgymCLI - Documented
gym listandgym searchextensions, external plugin discovery, and MCP auto-exposure - Added
gym eval reverifyand multi-reward verification contract documentation
Release Assets
v0.4.0
Release Summary
NeMo Gym v0.4.0 expands evaluation tooling and agent integrations. It establishes a new monthly release cadence; we will continue to provide day-zero support for Nemotron models, datasets, and environments.
Highlights:
- Unified
gymCLI: find agents and benchmarks by name withgym list, and catch config mistakes early withgym env validate - Diagnose evaluations with BLADE, an analysis skill for agents that reads your evaluation results and produces an evidence-backed report of which tasks failed, why, and the highest-impact fix (e.g. to the agent harness, training, verifier, or prompt)
- Measure the impact of agent skills: run the same tasks with different skill sets and compare how each changes agent performance
- Run agents in isolated sandboxes through a new pluggable provider framework
- More agent harnesses out of the box, including OpenClaw, Pi, and OpenCode
- Connect to hosted inference providers: Fireworks, Together.ai, OpenRouter, and more
- New benchmarks across science, long-context, and interactive tasks
First-Time Contributors
We welcomed 20+ new contributors to this release! A few highlights:
- @marta-sd and @wprazuch led the CLI refactor and clearer config errors
- @hemildesai added the pluggable sandbox provider infrastructure and OpenSandbox as the first built-in
- @adil-a laid the groundwork for Gym-owned MCP resources servers, letting a server expose its tools over MCP
- @eric-tramel added the BunsenChem chemistry benchmark
- @jeffwillette added the long machine translation datasets and servers
Thank you to all the new contributors for helping make NeMo Gym better!
Command Line Interface
- One
gymcommand for the full workflow, withgym env,gym eval,gym list, andgym datasetsubcommands - Reference agents, benchmarks, and environments by name: use
gym listto see what is available gym env validatechecks your config for missing, malformed, or empty values before a run and reports actionable errors
Evaluation & Diagnostics
- Skill evaluation: measure how agent skills affect performance by running the same tasks with different skill sets. Skills apply at rollout time as a run-level knob, so one dataset works across all skill variants and every rollout is tagged for comparison
- BLADE (Benchmark Level Analysis and Diagnostics Engine): a built-in analysis skill that reads an agent run’s rollouts, metrics, and configs and produces an evidence-backed report of which tasks failed, why, and the highest-impact fix (e.g. harness, training, verifier, or prompt)
Sandboxing
- Run tool-using and coding agents in isolated sandboxes through a pluggable provider framework
- Built-in OpenSandbox and Apptainer providers, with third-party providers discoverable via entry points
Configure Agent Harnesses
New harnesses join the existing built-in set (Claude Code, Hermes, OpenHands, and more):
- Added OpenCode, OpenClaw, and Pi agents for evaluation
- Claude Code runtime capabilities (tool access, MCP servers, and bare vs. native auto-discovery mode) are now easily set via the server config
Configure Models
- New
inference_providermodel server connects to any OpenAI-compatible hosted provider (Fireworks, Together.ai, OpenRouter, DeepInfra, Gemini, and more) with ready-made configs - Every Gym model server now speaks the Anthropic Messages API, so Anthropic-native harnesses like the Claude Code CLI can run against any model you serve with Gym — see Anthropic Messages
New Benchmarks
- Science: CritPt (research-level physics), SciCode (scientific coding), BunsenChem (chemistry multiple-choice), and FrontierScience Research (rubric-scored science)
- Long context: Graphwalks (long-context graph reasoning) and Long Machine Translation (PG19, WMT24++)
- Interactive: TALES, a text-adventure game suite
See the Available Environments table for the full list.
Deprecation Notices
- The legacy
ng_*andnemo_gym_*CLI commands (such asng_runandng_collect_rollouts) are deprecated in favor of the unifiedgymCLI. They still work for now but will be removed in a future release.
Bug Fixes
- Fixed intermittent connection errors during high-concurrency rollout collection
- Clear error messages instead of crashes when a config file contains invalid YAML
Documentation
- New Build Verifiers section with verification patterns and multi-reward verification
- New Evaluate section covering benchmarks, evaluation metrics, and a guide to agent-native results diagnostics
- New page for configuring and evaluating agent skills
Release Assets
v0.3.0
Release Summary
NeMo Gym v0.3.0 ships alongside the NVIDIA Nemotron 3 Ultra model release, open sourcing the environments and corresponding datasets used during training.
Highlights:
- 70+ new environments, including benchmarks such as Tau2 and Nemotron RL training environments
- Popular harness available out-of-the-box such as Claude Code and Hermes
- Integrations with OpenEnv and Harbor - use environments from these libraries directly with NeMo Gym
- Integration with VeRL - train with VeRL and scale rollout collection with NeMo Gym
First-Time Contributors
We welcomed 30+ new contributors to this release! Here are a few highlights:
- @grace-lam added the integration to run Harbor environments with NeMo Gym
- @aleksficek — added Competitive Coding Challenges environment
- @jthomson04 improved rollout resilience when models emit malformed tool-call arguments or missing message content
Thank you to all the new contributors for helping make NeMo Gym better!
New Environments & Benchmarks
Added 70+ new environments including novel datasets and integrations of popular benchmarks. New coverage spans:
- Coding — competitive programming, code infilling, SQL generation, and software-engineering benchmarks with execution-based verification
- Math & proofs — olympiad-style problems, proof grading and validation, and formal verification (including Lean)
- Knowledge & science — graduate-level QA, chemistry and physics tasks, and lab-style reasoning (including multimodal figure, table, and protocol tasks)
- Agentic — multi-turn tool use, search, sandboxed execution, finance workflows, and tau-bench-style conversational agents
- Instruction following — format constraints, citation compliance, and IFBench-style rule verification
- Safety & RLHF — jailbreak detection, abstention calibration, prompt-injection resistance, and generative reward modeling
- Multimodal, speech & translation — VLM benchmarks, visual grounding, ASR evaluation, and machine-translation quality metrics
- Chat & broad knowledge — arena-style preference evaluation and MMLU-family benchmarks
- Interactive RL — Gymnasium-style multi-step environments for spatial and game-based training
See the Available Environments table for the full list.
Configure Agent Harnesses
- Claude Code — available out of the box in NeMo Gym
- Hermes — available out of the box in NeMo Gym
- LangGraph agent — an adapter that lets you build custom agents using LangGraph patterns (reflection, subagent orchestration, parallel thinking, rewoo)
- Gymnasium agent — generic multi-turn harness for use with OpenAI Gym-style environments
Configure Models
- Optional
max_concurrent_requestson the OpenAI model server to cap in-flight API calls — useful for rate-limited external endpoints when rollout concurrency is high
Rollout Collection & Profiling
- New
ng_aggregate_rolloutscommand to merge rollout shards collected independently across multiple nodes, enabling distributed eval without requiring a single coordinated collection job
Environment Library Integrations
- OpenEnv — combine OpenEnv environments with NeMo Gym environments
- Harbor — combine Harbor environments with NeMo Gym environments
Deprecation Notices
- Documentation has moved from Sphinx to Fern. Old Sphinx URLs redirect to the new site at docs.nvidia.com/nemo/gym. The retired
docs/directory has been removed.
Bug Fixes
- Fixed aiohttp connection limit exhaustion under FastAPI/Uvicorn with multiple workers
- Fixed session cookie propagation for Starlette >= 1.0.0
- Fixed duplicated usage counting and errors on empty usage in subsequent model calls
- Improved rollout resilience when models emit malformed tool-call arguments or missing message content
- Fixed prompt-key hashing when inputs contain Pydantic BaseModel objects
Documentation
- New concepts pages for environments, evaluation, and training
- Improved Architecture page to clarify how environments map to NeMo Gym components
- Consolidated detailed setup and quickstart into a single improved quickstart with clearer descriptions
- Expanded Ecosystem page with environment library, training framework, and agent harness integrations
Release Assets
v0.2.1
Fixed PyPI package distribution that was broken in v0.2.0. No functional changes — all features and fixes from v0.2.0 apply.
v0.2.0
NeMo Gym v0.2.0 ships alongside the NVIDIA Nemotron 3 Super model release, open sourcing the RL environments and corresponding datasets used during training. This release adds 17 new training environments across coding, math, science, reasoning, agentic tasks, and safety, plus integrations with Aviary, Reasoning Gym, and Verifiers to combine additional environments. You can now run end-to-end rollout collection locally with vLLM and install directly from PyPI.
New Environments
Added 17 new resources servers spanning:
- Coding: Text to SQL, SWE RL Gen, SWE RL LLM Judge
- Math: Lean4 Mathematical Proofs
- Science: Aviary, NewtonBench
- Reasoning: MultiChallenge, ARC-AGI
- Agent tasks: xLAM Function Calling, Tavily Search, Single Step Tool Use, Terminus Judge, NeMo Skills Tools
- Safety: Jailbreak Detection, Over Refusal Detection
- RLHF: Generative Reward Model Compare
Added 5 new agent servers: Aviary agent, proof refinement agent, SWE agents, tool simulation agent, and verifiers agent.
Environment library integrations: Future House Aviary, Open-Thought Reasoning Gym, Prime Intellect Verifiers.
Model Serving
- Local vLLM model server with end-to-end rollout collection without an external API
- vLLM 0.16+ support for the reasoning field in responses
- Per-task chat templates and extra body args to support different model configurations across environments in multi-environment training
Rollout Collection & Profiling
- New
ng_reward_profilecommand to compute per-task pass rates and aggregate metrics - CPU profiling for rollout performance analysis
- Seeding on num_repeats for reproducible rollouts
Infrastructure & Developer Experience
- PyPI compatibility: install via
pip install nemo-gym - Dry run mode:
ng_run +dry_run=trueto validate configs and install environments without starting servers ng_statuscommand to list running servers and their health- FastAPI worker support for higher throughput across multiple workers
- Server stdout/stderr redirection with server name prefixes
Model Recipes
- Nemotron 3 Nano 30B end-to-end training recipe with single-GPU and multi-node tutorials
Documentation
- Added training tutorials for Unsloth, TRL, and Nemotron 3 Nano (single-GPU and multi-node)
- Added environment tutorials for creating environments, custom data preparation, and integrating external libraries
- Rewrote concepts documentation with new training approaches page, architecture diagrams, and expanded agent/resources server docs
- Revamped ecosystem page with training framework and environment library integrations
- Added deployment topology and SWE RL infrastructure case study
- Site-wide quality sweep: consistent naming, style guide, redirects, and FAQ additions
Bug Fixes
- Fixed 0.1.1 environments to work correctly with RL training pipelines
- Fixed crash when server receives malformed JSON during rollout collection
- Fixed dry run mode failing after initial implementation
- Fixed nested
responses_create_paramsoverrides not merging correctly from CLI - Fixed
ng_prepare_datafailing when multiple environments define overlapping metrics - Fixed reward profiling failing when model response doesn’t include usage stats
- Fixed NeMo-Skills python tool to use HTTP calls instead of subprocess execution
- Bumped Pillow and other packages to address security vulnerabilities
ng_dump_confignow redacts API key values from output
First-Time Contributors
We’d like to highlight the following first-time contributors:
- @sidnarayanan added the Aviary integration to enable training on any Aviary environment, a library of interactive RL environments spanning math, science, biology, and more
- @3mei added the text-to-SQL environment to generate SQL queries from natural language across multiple SQL dialects
- @Kelvin0110 added the NewtonBench environment to discover scientific laws through interactive experimentation
v0.1.1
Initial public release of NeMo Gym.