> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# Key Terminology

> Essential vocabulary for agent evaluation, policy-model training, RL workflows, and NeMo Gym.

Essential vocabulary for agent evaluation, policy-model training, RL workflows, and NeMo Gym. You'll encounter these terms throughout the tutorials and documentation.

Deeper treatment: [Environments](/about/concepts/environments), [Evaluation](/about/concepts/evaluation), [Training](/about/concepts/training), [Architecture](/about/architecture).

## Overview

A **dataset** is a JSONL file of **tasks** (one row = one problem). Running the agent on a task produces a **rollout** (also called a **trajectory**) — one attempt and its record.

Each run follows:

1. **seed\_session** — set up isolated state for this attempt
2. **agent loop** — call the model (policy), use tools, repeat until done
3. **verify** — score the attempt → **reward**

Two compositions matter:

|                 | Built from                                           | Notes                                                                                             |
| --------------- | ---------------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| **Agent**       | **Model** + **Agent harness**                        | The harness owns the loop, tools, context, and stop conditions.                                   |
| **Environment** | **Dataset** + **Verifier** + **State** (+ env tools) | The task world the agent acts on and is scored against. The **model is outside** the environment. |

The **agent harness** belongs to the agent, not the environment. In Gym packaging, an environment *config* still names which agent server to run — that is a wiring reference, not ownership of the harness.

The **model** (also called the **policy**) is what you call for generation — local weights or a remote API endpoint. In Gym it is usually exposed by the **Model server** (or your harness can call an endpoint directly). The same environment shape can be used for **benchmark** eval (fixed taskset / protocol) or for **training** (rewards or synthetic data); the main differences are task split and contamination controls, not a different environment type.

| Concept                                | In NeMo Gym                                                                                                                                                                              |
| -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Dataset                                | JSONL rows with `responses_create_params` + `verifier_metadata` — [Prepare Data](/data)                                                                                                  |
| Agent harness                          | Implemented by the **Agent server** (`responses_api_agents/`) — [Agent Server](/agent-server)                                                                                            |
| Verifier, per-attempt state, env tools | Implemented by the **Resources server** (`resources_servers/`): `seed_session()` initializes state; tools mutate it; `verify()` scores the attempt — [Build Verifiers](/build-verifiers) |
| Model                                  | Served by the **Model server** (`responses_api_models/`) — wraps local or remote endpoints — [Model Server](/model-server)                                                               |
| Protocol                               | **Responses API** (Chat Completions converted via middleware when needed)                                                                                                                |

## Glossary

### Tasks & rollouts

| Term                     | Definition                                                                                                                                                                                                                                                                                                                                                                                                             |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Task**                 | One problem for the agent to solve — one dataset row (`responses_create_params` + `verifier_metadata`).                                                                                                                                                                                                                                                                                                                |
| **Dataset**              | Collection of tasks plus metadata needed for scoring (usually JSONL). See [Prepare Data](/data).                                                                                                                                                                                                                                                                                                                       |
| **Rollout / trajectory** | **Rollout** (verb): execute an agent in an environment — take actions and record what happens. **Rollout** (noun) / **trajectory**: the ordered record of one attempt (states, actions, rewards). In Gym these names are aliases; architecture calls the resulting trajectory a *rollout*. Multiple rollouts per task support metrics like pass\@k. See [trajectory capabilities](/reference/trajectory-capabilities). |
| **Task attempt**         | One rollout for a specific task. Multiple attempts per task capture different approaches and support pass\@k.                                                                                                                                                                                                                                                                                                          |
| **Trace**                | Debug-oriented log of a rollout (timing, tool I/O, metadata) beyond the scored trajectory.                                                                                                                                                                                                                                                                                                                             |
| **Rollout batch**        | Multiple rollouts generated together (across tasks for throughput, or grouped on one task for methods like GRPO).                                                                                                                                                                                                                                                                                                      |
| **Rollout collection**   | Running inference, tools, and verification at scale to produce scored rollouts. Start with the [Quickstart](/get-started/quickstart) or [Evaluate](/evaluation).                                                                                                                                                                                                                                                       |

### Environment

| Term            | Definition                                                                                                                                                                                                                                                                          |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Environment** | The task world an agent runs against, excluding the model: `Dataset + Verifier + State` (plus environment tools). Gym configs also reference an **agent harness** for packaging. See [Environments](/about/concepts/environments) and [Build Environments](/environment-tutorials). |
| **State**       | Per-attempt mutable world (files, DB, tool results, and so on). `seed_session()` starts a clean session; tools update that session; it is not a snapshot of the whole environment config.                                                                                           |
| **Verifier**    | Scores a task attempt into a **reward** (typically 0–1) via `verify()` on the resources server. Also called scorer or grader. See [Build Verifiers](/build-verifiers).                                                                                                              |
| **Reward**      | Numerical score (typically 0.0–1.0) for how well the attempt did — used as eval metrics and as the RL learning signal.                                                                                                                                                              |
| **Sandbox**     | Runtime with isolated execution per attempt (e.g. one container). Broader runtimes (local process, Docker, Apptainer) host execution. See [Sandboxes](/infrastructure/sandbox).                                                                                                     |

### Agent & model

| Term                     | Definition                                                                                                                                                                                   |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Model / policy model** | The model being trained or evaluated — local weights or a remote endpoint, usually via the Model server. See [Model Server](/model-server).                                                  |
| **Agent harness**        | How the model interacts with the environment: loops model calls, routes tools, manages context, and decides when the task is done. See [Agent Server](/agent-server).                        |
| **Agent**                | Model + agent harness.                                                                                                                                                                       |
| **Multi-turn**           | Dialogue across turns where conversation context (and often state) persists.                                                                                                                 |
| **Multi-step**           | Sequential tool calls or intermediate steps in the agent loop before completion (may be single- or multi-turn). See [Multi-Step Environment](/environment-tutorials/multi-step-environment). |
| **Tool use**             | Function calling — invoking external capabilities (APIs, code execution, databases, and so on).                                                                                              |

### Evaluation

| Term           | Definition                                                                                                                                                                              |
| -------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Evaluation** | Run an agent on tasks, score results, and measure performance. See [Evaluation](/about/concepts/evaluation) and the [Evaluate](/evaluation) workflow.                                   |
| **Benchmark**  | Repeatable evaluation built on an environment: fixed dataset, metrics, and comparison protocol. Not every environment is used as a benchmark. See [Benchmarks](/evaluation/benchmarks). |
| **pass\@k**    | Fraction of tasks with at least one success among *k* rollouts. See [Aggregate Metrics](/evaluation/aggregate-metrics).                                                                 |

### Training

| Term                                          | Definition                                                                                                                                                                                                                                 |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **SFT (Supervised Fine-Tuning)**              | Train from examples of good behavior (**demonstration data**: successful / high-reward rollouts).                                                                                                                                          |
| **RL (Reinforcement Learning)**               | Improve the policy through environment interaction and reward signals.                                                                                                                                                                     |
| **Online / offline**                          | **Online**: update the policy from rewards while interacting (e.g. GRPO). **Offline**: train from pre-collected rollouts (e.g. SFT, DPO). See [Offline training with rollouts](/tutorials/training-tutorials/offline-training-w-rollouts). |
| **DPO (Direct Preference Optimization)**      | Offline preference training from pairs of rollouts (preferred vs dispreferred).                                                                                                                                                            |
| **GRPO (Group Relative Policy Optimization)** | Online RL that compares groups of rollouts on the same task relative to each other. See [Training Tutorials](/tutorials/training-tutorials) (e.g. [NeMo RL GRPO](/tutorials/training-tutorials/nemo-rl-grpo)).                             |
| **Demonstration data**                        | SFT examples from successful (high-reward) rollouts.                                                                                                                                                                                       |
| **Preference pairs**                          | DPO data: same task, high-reward vs low-reward rollouts (preferred vs dispreferred).                                                                                                                                                       |

Responses API and Chat Completions are listed in the mapping table above; see [Architecture](/about/architecture) and [Model Server](/model-server).