New Environment

View as Markdown

An environment includes the dataset, agent harness, verifier, and state. The model is external to the environment. A Gym config connects these components and supplies the runtime wiring.

The onboarding flow below covers a complete environment or benchmark. A resources server may be a reusable scoring, tool, or state component, or it may ship a runnable composition and data. Manifest-backed onboarding currently writes canonical workloads under environments/ and benchmarks/; runnable configs still colocated under resources_servers/ remain visible as no-manifest migration entries.

For a guide to building your first resources server, refer to Single Step Environment.

Environment Manifest

Each newly onboarded environment or benchmark has a manifest.yaml. The manifest gives contributors, reviewers, and tooling one versioned description of authored metadata and mirrored composition without replacing Gym configuration. Gym defines this contract as a Pydantic model and can emit JSON Schema from it when another tool needs a language-neutral representation.

The manifest is the place to read this declaration, but not every field is authored there:

Field groupAuthority
Identity and catalog metadata: name, version, experimental, kind, integration_profile, domain, description, modality, licensing, and authorsManifest
Behavioral declarations: reward, determinism, and optional session, state, sandbox, provenance, and lifecycle fieldsManifest
Composition: resources_server, agent_server, model_server, datasets, rollout_driver, and grading_modeGym config; mirrored read-only in the manifest
Benchmark protocol: canonical_split and standard_prompt_configManifest

When composition changes, edit the Gym config and update its manifest mirror. The runtime continues to resolve and execute the config; it does not use the manifest as a second wiring system. version follows Semantic Versioning and is intended to identify the resolved composition rather than the manifest text alone. Validation reports the version but does not yet enforce immutable composition or require version bumps.

The repository’s reference environment is example_single_tool_call:

name: example_single_tool_call
version: 0.1.0
experimental: false
kind: environment
integration_profile: custom-gym-verifier
domain: agent
description: Single-step environment that rewards invoking the get_weather tool.
modality: text
licensing: Apache-2.0
authors:
- fsiino-nvidia
- bxyu-nvidia
reward:
range: [0.0, 1.0]
higher_is_better: true
determinism: unknown
resources_server: example_single_tool_call
agent_server: simple_agent
datasets:
- name: example
type: example
jsonl_fpath: resources_servers/example_single_tool_call/data/example.jsonl
num_repeats: 1
model_server: policy_model
lifecycle: active

New manifests are experimental: true by default. The NeMo Gym team can manually set this field to false. This flag does not imply certificate-backed validation, which applies to the specific workload version and configuration covered by a certificate. Unknown keys are validation errors at every manifest level.

Field Reference

FieldTypeRequirement and allowed values
namestringRequired; non-empty and must match the catalog path
versionstringRequired; non-empty, with Semantic Versioning recommended
experimentalbooleanOptional; defaults to true
kindstringRequired; environment or benchmark
integration_profilestringRequired; one of the profiles below
domainstringRequired; math, coding, agent, knowledge, instruction_following, long_context, safety, games, translation, e2e, rlhf, or other
descriptionstringRequired; non-empty
modalitystringRequired; non-empty
licensingstringOptional; defaults to unknown; SPDX expression, internal, proprietary, or unknown
authorslist of stringsRequired; non-empty and unique
rewardobjectRequired; range is two finite increasing numbers and higher_is_better is boolean
determinismstringOptional; seeded, stochastic, or unknown (default)
resources_serverstringRequired; mirrored from config
agent_serverstringRequired; mirrored from config
datasetslist of objectsRequired and non-empty; dataset names must be unique
model_serverstring or nullRequired for custom-gym-verifier and custom-gym-agent-loop
rollout_driverstring or nullRequired only for external-rollout-driver; Python module:function form
grading_modestring or nullOptional resources-server-specific selector
session_modelstring or nullOptional; episode or step
statestring or nullOptional; none or per_session
sandboxstring or nullOptional sandbox requirement
canonical_splitstring or nullRequired for benchmarks
standard_prompt_configstring or nullRequired for benchmarks
adopted_fromobject or nullOptional; requires source URL, upstream ref, and reconciliation date
lifecyclestringOptional; active (default) or deprecated

Each datasets item requires name, type, and jsonl_fpath. type is train, validation, example, or benchmark; benchmark datasets also require prepare_script. prompt_config is optional, and num_repeats defaults to 1.

Integration Profiles

integration_profile classifies where the episode is driven. custom-gym-verifier is the default path; the other profiles describe three existing extension points.

ProfileUse whenAdditional requirementGenerated extension point
custom-gym-verifierGym drives the standard agent loop and only the model’s answers are measuredmodel_serverNone
custom-gym-agent-loopAuthored agent behavior, such as browsing or planning strategy, is part of the measurementmodel_serverAgent responses()
external-agent-loopAn external framework owns the episode and adapts its result back to GymNoneAgent responses() and run()
external-rollout-driverRollout coordination runs above the agentrollout_driverRollout-driver module

Every profile also requires a resources_server, an agent_server, and at least one dataset. rollout_driver is valid only for external-rollout-driver. A benchmark additionally requires canonical_split, standard_prompt_config, and at least one benchmark dataset with prepare_script; its dataset-level prompt_config remains optional.

The profile is an authored classification, not a runtime dispatcher. Validation recognizes the default Gym agent loop, custom Gym agent behavior, external agent adapters, and external rollout drivers. It reports unknown when static inspection is inconclusive and warns when the declaration disagrees. Runtime behavior remains config-driven; concrete component versions, capabilities, pinning, and swap constraints are not yet enforced.

Validation Layers

Environment validation is progressive. Manifest conformance checks the workload declaration. Static manifest validation checks whether that declaration matches the repository configuration. Verifier fixtures exercise representative scoring behavior, while evaluation runs and reward profiling establish runtime behavior and grading quality.

Static manifest validation runs without starting the workload. It:

  • validates the manifest schema and workload identity;
  • resolves Gym configuration and inheritance without runtime side effects;
  • compares manifest composition mirrors with the authoritative configuration;
  • reports an inferred integration profile and the static evidence behind it;
  • parses referenced preparation and rollout-driver hooks without importing or executing them; and
  • streams declared JSONL data, applies benchmark prompts row by row, and checks the resulting rollout inputs.

With synchronization enabled, it writes corrected composition mirrors atomically only after every static check passes. Authored metadata remains unchanged.

It does not import components, start services, execute preparation code, call a model, run an evaluation, probe model or resources endpoints, or check evaluation-output writability. Runtime pre-run readiness is a separate validation layer. Component versions and capabilities, profile pinning, and adopted_from source resolution are deferred from this initial static validation. Scorer behavior and grading quality require the later validation layers.

Onboarding Commands

Search the local catalog before creating a workload. The onboarding journey then uses four commands to scaffold, validate, test, and publish it:

gym search "task description"
gym env init --environment my_eval --profile custom-gym-verifier
gym env validate my_eval
gym env test my_eval
gym env publish my_eval

The scaffold creates a manifest, Gym config, sample data, and README. Names that create Python components must be lowercase Python identifiers. Benchmarks also receive source data, a prompt, and a prepare() function. A new scorer adds a resources server and verifier fixture. Non-default templates expose the selected extension point but initially delegate to existing Gym behavior; replace their generated TODO before review. For an external agent loop, replace run() with the framework adapter and make responses() raise NotImplementedError.

To reuse a scorer that already exports the verifier-fixture contract, declare its reward contract instead of copying its implementation:

gym env init --benchmark my_benchmark --profile custom-gym-verifier \
--reuse-verifier existing_scorer --reward-range 0 1 --higher-is-better

Scaffolding is non-destructive: an identical rerun is a no-op, and any conflicting file aborts the complete write set. gym env validate --sync NAME updates only mirrored composition fields after all static checks pass. gym env test --update-expected NAME updates fixture rewards only after every behavioral check passes. gym env publish NAME runs validation and the fixture, rejects manifest metadata placeholders, and confirms that the exact manifest is discoverable. The manifest is the registry record, so this structural check is idempotent and does not commit or push changes.

gym list environments and unqualified gym search read manifests and legacy runnable configs together. Entries with experimental: true carry an experimental annotation; entries with experimental: false have no status annotation. Unmigrated entries are labeled no-manifest. Reusable resources-server components without agent composition and datasets remain available through gym list resources-servers and are not environment entries.

Publication includes an experimental annotation only when the manifest’s flag is true. CODEOWNERS updates, immutable version enforcement, capability checks, certificate-backed validation, and a hosted catalog index are not yet automated.

Verifier Fixture Contract

The resources server owns and exports one VERIFIER_FIXTURE, so every workload that reuses the scorer also reuses its scoring tests. The fixture requires three cases, plus a fourth when the manifest declares seeded:

  • a full-reward case that reaches the better endpoint declared by higher_is_better;
  • a zero-reward case that reaches the opposite endpoint;
  • a malformed request that fails as declared; and
  • for a seeded environment, the same request producing the same reward after an explicit reseed on fresh server instances.

Fixture execution runs directly in the resources server’s dependency environment and does not start Gym services or Ray. The first run prepares that environment in the same way as existing server tests. Updating expected rewards is explicit and atomic; range, endpoint, malformed-input, and determinism checks still apply. gym env init --reuse-verifier checks that the selected resources-server entrypoint declares a fixture, and gym env test executes it. A shared fixture attests the scorer itself, so a workload that overrides grading_mode needs workload-specific cases.

Guiding Principles

Adding a training environment has the same local correctness requirements as Adding A Benchmark: its manifest must validate and its verifier fixture must pass. Behavior-changing environment or agent work must also run representative real smoke rollouts as required by AGENTS.md. A full evaluation, reward profile, or training run can provide stronger evidence about measurement quality and training utility, but it is optional and is not a publication or merge compute gate.

When compute is available, a useful training experiment isolates the environment’s effect on the targeted capability. GRPO with NeMo RL, 64 prompts per step, and 16 rollouts per prompt is one starting point; adjust it to the environment and available compute.

If you run this experiment, use a model that achieves meaningful performance during reward profiling and include the relevant configuration, curves, and links in the pull request.

Required Files

Your resources server must include these files:

FileDescription
app.pyMain server implementation with verify function
configs/*.yamlConfiguration with valid domain field
tests/test_app.pyAt least one unit test
Example dataAt least one representative input at a path referenced by the workload config
requirements.txtPython dependencies
README.mdDocumentation with licensing information

Optional rollout evidence may be saved in data/example_rollouts.jsonl.

Contribution Workflow

Contributing a resources server follows this sequence:

StepPhaseDescription
1Curate TasksCollect or generate training tasks and create example data
2ImplementationBuild resources server with verification logic
3TestingWrite and run unit tests
4Smoke RolloutsFor behavior changes, run a representative model rollout and inspect agent and verifier behavior
5Reward ProfilingOptionally inspect reward distribution for training environments; benchmarks require profiling
6Training ValidationOptionally test training utility
7Submit PRSubmit pull request with all required information
8ReviewAddress feedback on the required local checks and contribution metadata

Detailed Steps

1. Curate Training Tasks

Prepare the dataset for your environment:

  • Collect or generate prompts/tasks for your environment
  • Create data/example.jsonl with at least one representative task example

2. Resources Server Implementation

Build your resources server:

  • Run gym env init --resources-server my_server to scaffold the new resources server
  • Follow the Single Step Environment guide to implement your specific logic
  • Implement verification logic for your tasks by defining the verify() function
  • Set the domain field in your resources server configuration (see Domain).
  • Complete the auto-generated README.md with licensing information

3. Testing

Write and run tests for your resources server:

  • At least one test per server is required for PR approval
  • You are responsible for ensuring your tests adequately cover your server’s functionality

4. Generate Example Rollouts (Required for Behavior Changes)

When environment or agent work changes runtime behavior, configure a model endpoint, run a representative rollout, and inspect the agent and verifier results. For example:

gym env start \
--resources-server my_server \
--model-type openai_model
gym eval run --no-serve \
--agent your_agent \
--input path/to/example.jsonl \
--output results/my_server_smoke.jsonl \
--limit 1

Document the commands and observed behavior in the PR. For a manifest-backed workload, saving the output under data/example_rollouts.jsonl is optional. A standalone legacy resources-server contribution must retain the five-row data/example_rollouts.jsonl artifact required by its data validator until that validator is migrated. Metadata-only catalog or manifest changes and docs-only changes do not require model compute.

5. Reward Profiling (Optional for Training Environments)

Fixed evaluation benchmarks must follow the reward-profiling requirements in Adding A Benchmark. For training environments, profiling remains optional:

Run inference to inspect reward distribution:

  • Use a ~500 sample subset (minimum)
  • Use Qwen3-4B, Qwen3 30B A3B, or equivalent model
  • Generate 16 responses per prompt
  • Report reward distribution
  • For tool calling: Provide tool call metrics and correlation with rewards

6. Training-Based Validation (Optional)

Validate with actual training:

  • Train with GRPO on Qwen3-4B, Qwen 30B A3B Instruct, or equivalent model
  • Include training accuracy curve
  • Include test benchmark accuracy curve (if applicable)

7. Submit PR

Include the following in your pull request description:

  • Description of the environment
  • Description of the verification logic
  • Description of the prompts/tasks: What is the source? Which domain does it cover?
  • Provide relevant license information for data and software. If models were used for synthetic data generation, note this in your PR description

8. PR Review Process

After submitting your PR:

  1. A team member reviews the manifest, composition, fixture, and licensing information
  2. Address any feedback from reviewers
  3. After approval, maintainers merge the contribution

Reviewers inspect the required smoke-rollout evidence for behavior-changing work and may inspect optional full evaluation or training evidence when provided. They do not need to reproduce a compute-heavy run for the contribution to merge.

For optimal performance and scalability, we recommend following these design patterns:

Async-First Design

Endpoint handlers should be asynchronous to handle concurrent requests efficiently during training:

# Recommended: async function
async def verify(self, body: BaseVerifyRequest) -> BaseVerifyResponse:
return BaseVerifyResponse(**body.model_dump(), reward=1.0)

Avoid spawning additional threads or processes unless necessary. A single Gym instance can handle tens of thousands of concurrent requests when properly implemented.

NeMo Gym OpenAI Client

We recommend using the NeMo Gym OpenAI client. Import it and the core types from the top-level nemo_gym package:

from nemo_gym import (
NeMoGymAsyncOpenAI,
NeMoGymResponse,
NeMoGymResponseCreateParamsNonStreaming,
)

The NeMo Gym client is optimized for scale and provides consistent behavior. External clients like LiteLLM often preprocess or postprocess inputs and outputs in ways that can interfere with training data collection.

Pydantic Models

Consider using Pydantic models for request and response validation by extending base classes imported from the top-level nemo_gym package:

from pydantic import BaseModel
from nemo_gym import BaseVerifyRequest, BaseVerifyResponse
class MyVerifyRequest(BaseVerifyRequest):
expected_result: str
difficulty: int

Error Handling

Tool execution errors should be propagated back to the model rather than crashing the server, enabling the model to learn from mistakes:

async def execute_tool(self, path: str, body: ToolRequest) -> ToolResponse:
try:
result = self.tool_functions[path](**body.model_dump())
return ToolResponse(output=result)
except Exception as e:
# Return error to model so it can correct itself
return ToolResponse(output=f"Error executing tool '{path}': {str(e)}")

Configuration

Pass configuration through NeMo Gym config files rather than environment variables for better reproducibility:

# configs/my_server.yaml
host: 0.0.0.0
port: 8000
domain: agent

Multi-Step Rollouts

For multi-step scenarios, the model returns training information on response messages (prompt_token_ids, generation_token_ids, generation_log_probs). When constructing messages for subsequent model calls, propagate this information from previous responses to maintain the training data chain.

Reference