New Environment

View as Markdown

New Environment

An environment is the logical execution composition of a task, agent harness, model bindings, resources, and episode protocol. A Gym config selects and connects these participants. The task defines what must be done and how it is scored; the harness decides how to solve it.

The onboarding flow below covers a complete environment or benchmark. A resources server may be a reusable scoring, tool, or state component, or it may ship a runnable composition and data. Manifest-backed onboarding currently writes canonical workloads under environments/ and benchmarks/; runnable configs still colocated under resources_servers/ remain visible as no-manifest migration entries.

For a guide to building your first resources server, refer to Single Step Environment.

Component Ownership

The agent/environment separation design keeps task behavior independent of the selected harness:

PartResponsibility
Task / Resources ServerTask setup, state, tools, verification, task resources such as sandboxes, and their cleanup
Agent ServerHarness behavior, model interaction, dependencies, and harness-local sessions
Model ServerModel inference and provider-specific adaptation
EnvironmentConfiguration that selects the task, participants, model bindings, episode protocol, and run policy
Environment ServerEpisode entrypoint and protocol orchestration; legacy adapters delegate episode execution to the agent

A Resources Server is one participant in an environment. Reuse an existing harness when adding tasks or scoring; a new benchmark does not by itself need a new Agent Server. Expose every tool or artifact the verifier depends on through the task contract. For example, a verifier that reads deliverables must document how any compatible harness can create them, without relying on a private handoff in one agent.

Every runnable agent must be referenced by an Environment Server in Gym config; rollout collection routes POST /run through that server. The scaffold currently wires legacy_agent, which forwards the agent’s existing /run and /aggregate_metrics contracts without taking over its lifecycle controls. Native protocols such as single_agent_turn coordinate participant sessions, verification, timeouts, and cleanup. Migrating legacy components to those protocols is separate from adding the compatibility routing. The manifest describes composition; it does not perform that migration or define an environment_server field.

An external framework may expose a reusable harness, task tools and scoring, or an entire episode engine. Adapt the reusable components where possible and document the boundary when the engine cannot be separated. Existing external-agent-loop adapters run behind the compatibility Environment Server; moving a fused engine to its own native episode protocol is a separate migration tracked by the composability epic.

Environment Manifest

Each newly onboarded environment or benchmark has a manifest.yaml. The manifest gives contributors, reviewers, and tooling one versioned description of authored metadata and mirrored composition without replacing Gym configuration. Gym defines this contract as a Pydantic model and can emit JSON Schema from it when another tool needs a language-neutral representation.

The manifest is the place to read this declaration, but not every field is authored there:

Field groupAuthority
Identity and catalog metadata: name, version, experimental, kind, integration_profile, domain, description, modality, licensing, and authorsManifest
Behavioral declarations: reward, determinism, and optional session, state, sandbox, provenance, and lifecycle fieldsManifest
Composition: resources_server, agent_server, model_server, datasets, rollout_driver, and grading_modeGym config; mirrored read-only in the manifest
Configuration selection: config_path and optional dataset_ownerManifest
Benchmark protocol: canonical_split, prompt_source, and standard_prompt_configManifest; consistent with preparation and runtime prompting

When composition changes, edit the Gym config and update its manifest mirror. The runtime continues to resolve and execute the config; it does not use the manifest as a second wiring system. version follows Semantic Versioning and is intended to identify the resolved composition rather than the manifest text alone. Validation reports the version but does not yet enforce immutable composition or require version bumps.

The repository’s reference environment is example_single_tool_call:

name: example_single_tool_call
version: 0.1.0
experimental: false
kind: environment
integration_profile: custom-gym-verifier
domain: agent
description: Single-step environment that rewards invoking the get_weather tool.
modality: text
licensing: Apache-2.0
authors:
- fsiino-nvidia
- bxyu-nvidia
reward:
range: [0.0, 1.0]
higher_is_better: true
determinism: unknown
resources_server: example_single_tool_call
agent_server: simple_agent
datasets:
- name: example
type: example
jsonl_fpath: resources_servers/example_single_tool_call/data/example.jsonl
num_repeats: 1
model_server: policy_model
lifecycle: active

New manifests are experimental: true by default. The NeMo Gym team can manually set this field to false. This flag does not imply certificate-backed validation, which applies to the specific workload version and configuration covered by a certificate. Unknown keys are validation errors at every manifest level.

Field Reference

FieldTypeRequirement and allowed values
namestringRequired; non-empty and must match the catalog path
versionstringRequired; non-empty, with Semantic Versioning recommended
experimentalbooleanOptional; defaults to true
kindstringRequired; environment or benchmark
integration_profilestringRequired; one of the profiles below
domainstringRequired; math, coding, agent, knowledge, instruction_following, long_context, safety, games, translation, e2e, rlhf, or other
descriptionstringRequired; non-empty
modalitystringRequired; non-empty
licensingstringOptional; defaults to unknown; SPDX expression, internal, proprietary, or unknown
authorslist of stringsRequired; non-empty and unique
rewardobjectRequired; range is two finite numbers ordered lower <= upper (a constant range is allowed) and higher_is_better is boolean
determinismstringOptional; seeded, stochastic, or unknown (default)
config_pathstringOptional; defaults to config.yaml beside the manifest; must be relative and stay inside the catalog tree
dataset_ownerstring or nullOptional; selects the dataset-bearing config instance when selection is ambiguous
resources_serverstring or nullRequired; mirrored from config; null is permitted except for custom-gym-verifier
agent_serverstringRequired; mirrored from config
datasetslist of objectsRequired and non-empty; dataset names must be unique
model_serverstring or nullRequired for custom-gym-verifier and custom-gym-agent-loop
rollout_driverstring or nullRequired only for external-rollout-driver; Python module:function form
grading_modestring or nullOptional resources-server-specific selector
session_modelstring or nullOptional; episode or step
statestring or nullOptional; none or per_session
sandboxstring or nullOptional sandbox requirement
canonical_splitstring or nullRequired for benchmarks
prompt_sourcestringOptional; template (default), prepared, or agent; see below
standard_prompt_configstring or nullRequired for template benchmarks; absent for prepared prompts; optional for agent-owned prompting
adopted_fromobject or nullOptional; requires source URL, upstream ref, and reconciliation date
lifecyclestringOptional; active (default) or deprecated

Each datasets item requires name, type, and jsonl_fpath. type is train, validation, example, or benchmark; benchmark datasets also require prepare_script. num_repeats defaults to 1. canonical_split records the source split being measured, while dataset type: benchmark selects its role in Gym.

For named configs, put the manifest in the workload’s catalog directory and set config_path relative to it. For example, benchmarks/my_benchmark/config_with_system/manifest.yaml can declare name: my_benchmark/config_with_system and config_path: ../config_with_system.yaml. Shared prompts and suite fragments are not standalone workloads.

dataset_owner names a config instance, not a component implementation. It selects the intended dataset declaration when multiple instances have datasets; runtime agent routing must still be unambiguous. New task scaffolds put datasets on the resources-server instance. Existing agent-owned declarations, including the reuse-verifier scaffold, remain supported.

Prompt Sources

prompt_sourceContract
template (default)Prepared rows contain template fields. For benchmarks, standard_prompt_config is required and any dataset prompt_config must match it.
preparedEach benchmark row already contains non-empty responses_create_params.input. Omit both standard_prompt_config and dataset prompt_config.
agentThe harness or external engine constructs prompts during execution. Omit dataset prompt_config; an optional standard_prompt_config documents an agent template and is not applied to dataset rows by validation.

Choose the source that matches the implemented prompting path. Changing this metadata alone does not change runtime prompting.

Integration Profiles

integration_profile classifies where the episode is driven. custom-gym-verifier is the default path; the other profiles describe three existing extension points.

ProfileUse whenAdditional requirementGenerated extension point
custom-gym-verifierGym drives the standard agent loop and only the model’s answers are measuredmodel_serverNone
custom-gym-agent-loopAuthored agent behavior, such as browsing or planning strategy, is part of the measurementmodel_serverAgent responses()
external-agent-loopAn external framework owns the episode and adapts its result back to GymNoneAgent responses() and run()
external-rollout-driverRollout coordination runs above the agentrollout_driverRollout-driver module

Every profile declares resources_server, agent_server, and at least one dataset. If the workload has no resources server, profiles other than custom-gym-verifier may explicitly declare resources_server: null. rollout_driver is valid only for external-rollout-driver. A benchmark additionally requires canonical_split and at least one benchmark dataset with prepare_script; its prompt requirements depend on prompt_source above.

The profile is an authored classification, not a runtime dispatcher. Validation recognizes the default Gym agent loop, custom Gym agent behavior, external agent adapters, and external rollout drivers. It reports unknown when static inspection is inconclusive and warns when the declaration disagrees. Runtime behavior remains config-driven; concrete component versions, capabilities, pinning, and swap constraints are not yet enforced.

Validation Layers

Environment validation is progressive. Manifest conformance checks the workload declaration. Static manifest validation checks whether that declaration matches the repository configuration. Verifier fixtures exercise representative scoring behavior, while evaluation runs and reward profiling establish runtime behavior and grading quality.

Static manifest validation runs without starting the workload. It:

  • validates the manifest schema and workload identity;
  • resolves Gym configuration and inheritance without runtime side effects;
  • compares manifest composition mirrors with the authoritative configuration;
  • reports an inferred integration profile and the static evidence behind it;
  • parses referenced preparation and rollout-driver hooks without importing or executing them; and
  • streams declared JSONL data, applies benchmark prompts row by row, and checks the resulting rollout inputs.

With synchronization enabled, it writes corrected composition mirrors atomically only after every static check passes. Authored metadata remains unchanged.

It does not import components, start services, execute preparation code, call a model, run an evaluation, probe model or resources endpoints, or check evaluation-output writability. Runtime pre-run readiness is a separate validation layer. Component versions and capabilities, profile pinning, and adopted_from source resolution are deferred from this initial static validation. Scorer behavior and grading quality require the later validation layers.

Onboarding Commands

Search the local catalog before creating a workload. The onboarding journey then uses four commands to scaffold, validate, test, and publish it:

gym search "task description"
gym env init --environment my_eval --profile custom-gym-verifier
gym env validate my_eval
gym env test my_eval
gym env publish my_eval

The scaffold creates a manifest, Gym config, sample data, and README. Names that create Python components must be lowercase Python identifiers. Benchmarks also receive source data, a prompt, and a prepare() function. A new scorer adds a resources server and verifier fixture. Non-default templates expose the selected extension point but initially delegate to existing Gym behavior; replace their generated TODO before review. An external-loop scaffold is an adapter starting point under the current runtime; document which responsibilities the framework owns and test the actual execution path.

To reuse a scorer that already exports the verifier-fixture contract, declare its reward contract instead of copying its implementation:

gym env init --benchmark my_benchmark --profile custom-gym-verifier \
--reuse-verifier existing_scorer --reward-range 0 1 --higher-is-better

Scaffolding is non-destructive: an identical rerun is a no-op, and any conflicting file aborts the complete write set. gym env validate --sync NAME updates only mirrored composition fields after all static checks pass. gym env test --update-expected NAME updates fixture rewards only after every behavioral check passes. gym env publish NAME runs validation and the fixture, rejects manifest metadata placeholders, and confirms that the exact manifest is discoverable. The manifest is the registry record, so this structural check is idempotent and does not commit or push changes.

Run gym eval prepare --benchmark NAME before validating a benchmark whose prepared data is not yet present. Static validation reads the declared JSONL but does not execute preparation code.

gym env test NAME and gym env publish NAME require a resources server exporting VERIFIER_FIXTURE. A manifest with resources_server: null can pass static validation, but these fixture commands do not support it. Use the external adapter’s own scoring tests and representative rollouts, and record the limitation in the PR. gym env test --resources-server NAME runs the component’s test suite rather than the workload fixture.

gym list environments and unqualified gym search read manifests and legacy runnable configs together. Entries with experimental: true carry an experimental annotation; entries with experimental: false have no status annotation. Unmigrated entries are labeled no-manifest. Reusable resources-server components without agent composition and datasets remain available through gym list resources-servers and are not environment entries.

Publication includes an experimental annotation only when the manifest’s flag is true. CODEOWNERS updates, immutable version enforcement, capability checks, certificate-backed validation, and a hosted catalog index are not yet automated.

Verifier Fixture Contract

The resources server owns and exports one VERIFIER_FIXTURE, so every workload that reuses the scorer also reuses its scoring tests. The fixture requires three cases, plus a fourth when the manifest declares seeded:

  • a full-reward case that reaches the better endpoint declared by higher_is_better;
  • a zero-reward case that reaches the opposite endpoint;
  • a malformed request that fails as declared; and
  • for a seeded environment, the same request producing the same reward after an explicit reseed on fresh server instances.

Fixture execution runs directly in the resources server’s dependency environment and does not start Gym services or Ray. The first run prepares that environment in the same way as existing server tests. Updating expected rewards is explicit and atomic; range, endpoint, malformed-input, and determinism checks still apply. gym env init --reuse-verifier checks that the selected resources-server entrypoint declares a fixture, and gym env test executes it. A shared fixture attests the scorer itself, so a workload that overrides grading_mode needs workload-specific cases.

Guiding Principles

Adding a training environment has the same local correctness requirements as Adding A Benchmark: its manifest must validate and its verifier fixture must pass where that contract is supported. Behavior-changing environment or agent work must also run representative real smoke rollouts as required by AGENTS.md. A full evaluation, reward profile, or training run can provide stronger evidence about measurement quality and training utility, but it is optional and is not a publication or merge compute gate.

When compute is available, a useful training experiment isolates the environment’s effect on the targeted capability. GRPO with NeMo RL, 64 prompts per step, and 16 rollouts per prompt is one starting point; adjust it to the environment and available compute.

If you run this experiment, use a model that achieves meaningful performance during reward profiling and include the relevant configuration, curves, and links in the pull request.

Required Files

Your resources server must include these files:

FileDescription
app.pyMain server implementation with verify function
configs/*.yamlConfiguration with valid domain field
tests/test_app.pyAt least one unit test
Example dataAt least one representative input at a path referenced by the workload config
requirements.txtPython dependencies
README.mdDocumentation with licensing information

Optional rollout evidence may be saved in data/example_rollouts.jsonl.

Contribution Workflow

Contributing a resources server follows this sequence:

StepPhaseDescription
1Curate TasksCollect or generate training tasks and create example data
2ImplementationBuild resources server with verification logic
3TestingWrite and run unit tests
4Smoke RolloutsFor behavior changes, run a representative model rollout and inspect agent and verifier behavior
5Reward ProfilingInspect reward distribution and variance when evaluating benchmark fidelity or training utility
6Training ValidationOptionally test training utility
7Submit PRSubmit pull request with all required information
8ReviewAddress feedback on the required local checks and contribution metadata

Detailed Steps

1. Curate Training Tasks

Prepare the dataset for your environment:

  • Collect or generate prompts/tasks for your environment
  • Create data/example.jsonl with at least one representative task example

2. Resources Server Implementation

Build your resources server:

  • Run gym env init --resources-server my_server to scaffold the new resources server
  • Follow the Single Step Environment guide to implement your specific logic
  • Implement verification logic for your tasks by defining the verify() function
  • Set the domain field in your resources server configuration (see Domain).
  • Complete the auto-generated README.md with licensing information

3. Testing

Write and run tests for your resources server:

  • At least one test per server is required for PR approval
  • You are responsible for ensuring your tests adequately cover your server’s functionality

4. Generate Example Rollouts (Required for Behavior Changes)

When environment or agent work changes runtime behavior, configure a model endpoint, run a representative rollout, and inspect the agent and verifier results. For example:

gym env start \
--resources-server my_server \
--model-type openai_model
gym eval run --no-serve \
--agent your_agent \
--input path/to/example.jsonl \
--output results/my_server_smoke.jsonl \
--limit 1

Document the commands and observed behavior in the PR. For a manifest-backed workload, saving the output under data/example_rollouts.jsonl is optional. A standalone legacy resources-server contribution must retain the five-row data/example_rollouts.jsonl artifact required by its data validator until that validator is migrated. Metadata-only catalog or manifest changes and docs-only changes do not require model compute.

5. Reward Profiling (Optional for Training Environments)

For benchmark fidelity checks, follow the reward-profiling guidance in Add a benchmark. For training environments, profiling remains optional:

Run inference to inspect reward distribution:

  • Use a ~500 sample subset (minimum)
  • Use Qwen3-4B, Qwen3 30B A3B, or equivalent model
  • Generate 16 responses per prompt
  • Report reward distribution
  • For tool calling: Provide tool call metrics and correlation with rewards

6. Training-Based Validation (Optional)

Validate with actual training:

  • Train with GRPO on Qwen3-4B, Qwen 30B A3B Instruct, or equivalent model
  • Include training accuracy curve
  • Include test benchmark accuracy curve (if applicable)

7. Submit PR

Include the following in your pull request description:

  • Description of the environment
  • Description of the verification logic
  • Description of the prompts/tasks: What is the source? Which domain does it cover?
  • Provide relevant license information for data and software. If models were used for synthetic data generation, note this in your PR description

8. PR Review Process

After submitting your PR:

  1. A team member reviews the manifest, composition, fixture, and licensing information
  2. Address any feedback from reviewers
  3. After approval, maintainers merge the contribution

Reviewers inspect the required smoke-rollout evidence for behavior-changing work and may inspect optional full evaluation or training evidence when provided. They do not need to reproduce a compute-heavy run for the contribution to merge.

For optimal performance and scalability, we recommend following these design patterns:

Async-First Design

Endpoint handlers should be asynchronous to handle concurrent requests efficiently during training:

# Recommended: async function
async def verify(self, body: BaseVerifyRequest) -> BaseVerifyResponse:
return BaseVerifyResponse(**body.model_dump(), reward=1.0)

Avoid spawning additional threads or processes unless necessary. A single Gym instance can handle tens of thousands of concurrent requests when properly implemented.

NeMo Gym OpenAI Client

We recommend using the NeMo Gym OpenAI client. Import it and the core types from the top-level nemo_gym package:

from nemo_gym import (
NeMoGymAsyncOpenAI,
NeMoGymResponse,
NeMoGymResponseCreateParamsNonStreaming,
)

The NeMo Gym client is optimized for scale and provides consistent behavior. External clients like LiteLLM often preprocess or postprocess inputs and outputs in ways that can interfere with training data collection.

Pydantic Models

Consider using Pydantic models for request and response validation by extending base classes imported from the top-level nemo_gym package:

from pydantic import BaseModel
from nemo_gym import BaseVerifyRequest, BaseVerifyResponse
class MyVerifyRequest(BaseVerifyRequest):
expected_result: str
difficulty: int

Error Handling

Tool execution errors should be propagated back to the model rather than crashing the server, enabling the model to learn from mistakes:

async def execute_tool(self, path: str, body: ToolRequest) -> ToolResponse:
try:
result = self.tool_functions[path](**body.model_dump())
return ToolResponse(output=result)
except Exception as e:
# Return error to model so it can correct itself
return ToolResponse(output=f"Error executing tool '{path}': {str(e)}")

Configuration

Pass configuration through NeMo Gym config files rather than environment variables for better reproducibility:

# configs/my_server.yaml
host: 0.0.0.0
port: 8000
domain: agent

Multi-Step Rollouts

For multi-step scenarios, the model returns training information on response messages (prompt_token_ids, generation_token_ids, generation_log_probs). When constructing messages for subsequent model calls, propagate this information from previous responses to maintain the training data chain.

Reference