For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
LogoLogoNeMo Gym
DocumentationAPI Reference
DocumentationAPI Reference
  • About
    • Concepts
    • Architecture
    • Ecosystem
    • Release Notes
  • Get Started
    • Prerequisites
    • Installation
    • Quickstart
  • Prepare Data
    • Prepare and Validate
    • Download from Hugging Face
    • Prompt Config
  • Configure Agents
    • Integrate Existing Agents
    • Drive a Remote Agent
    • Agent Skills
  • Configure Models
    • OpenAI
    • Azure OpenAI
    • Inference Providers
    • Model-call capture
    • vLLM
    • Local vLLM
    • Local vLLM Proxy
    • Anthropic Messages
  • Build Verifiers
    • Verification Patterns
    • Multi-Reward Verification
  • Evaluate
    • Benchmarks
    • Browse Environments
    • Aggregate Metrics
    • Diagnose Results
  • Tutorials
    • Training Tutorials
    • Evaluation Tutorials
      • Evaluate EvalPlus
      • BLADE Analysis Skill
      • Reverify Rollouts
  • Build Environments
    • Resources Server APIs
    • Single-Step Environment
    • Multi-Step Environment
    • Stateful Environment
    • MCP Resources Server
    • Real-World Environment
    • Integrate external libraries
  • Model Recipes
    • Nemotron 3 Nano
    • Nemotron 3 Super
  • Infrastructure
    • Deployment Topology
    • Sandbox API
    • Engineering Notes
  • Reference
    • Configuration
    • RL Framework Compatibility
    • CLI Commands
    • FAQ
    • Trajectory capability matrix
  • Troubleshooting
    • Configuration Errors
  • Contribute
    • Development Setup
    • Environments
    • Integrate RL Frameworks
    • Agent Skills
  • About
  • Concepts
  • Environments
  • Evaluation
  • Training
  • Key Terminology
  • Architecture
  • Ecosystem
  • Release Notes
  • Prerequisites
  • Installation
  • Quickstart
  • Prepare Data
  • Prepare and Validate
  • Download from Hugging Face
  • Prompt Config
  • Configure Agents
  • Integrate Existing Agents
  • Drive a Remote Agent
  • Agent Skills
  • Configure Models
  • OpenAI
  • Azure OpenAI
  • Inference Providers
  • Model-call capture
  • vLLM
  • Local vLLM
  • Local vLLM Proxy
  • Anthropic Messages
  • Build Verifiers
  • Verification Patterns
  • Equivalence Match
  • Execution and State Match
  • LLM-as-Judge
  • Multi-Reward Verification
  • Evaluate
  • Benchmarks
  • Browse Environments
  • Aggregate Metrics
  • Diagnose Results
  • Training Tutorials
  • NeMo RL
  • About Workplace Assistant
  • Gym Configuration
  • Multi-Node Training
  • NeMo RL Configuration
  • Setup
  • Single Node Training
  • Unsloth
  • Multi-Environment Training
  • Training with VeRL
  • Offline Training (SFT/DPO)
  • Evaluation Tutorials
  • Evaluate EvalPlus
  • BLADE Analysis Skill
  • Reverify Rollouts
  • Build Environments
  • Resources Server APIs
  • Single-Step Environment
  • Multi-Step Environment
  • Stateful Environment
  • MCP Resources Server
  • Real-World Environment
  • Generating Training Data
  • Resources Server Implementation
  • Integrate external libraries
  • Model Recipes
  • Nemotron 3 Nano
  • Nemotron 3 Super
  • Infrastructure
  • Deployment Topology
  • Sandbox API
  • OpenSandbox Provider
  • Apptainer Provider
  • Docker Provider
  • ECS Fargate
  • Adding a Sandbox Provider
  • Engineering Notes
  • aiohttp vs httpx
  • Responses API
  • SWE RL Case Study
  • System Design
  • Reference
  • Configuration
  • RL Framework Compatibility
  • CLI Commands
  • FAQ
  • Trajectory capability matrix
  • Troubleshooting
  • Configuration Errors
  • Contribute
  • Development Setup
  • Environments
  • Add a benchmark
  • New Environment
  • Integrate RL Frameworks
  • Generation Backend
  • Integration Footprint
  • On-Policy Corrections
  • Success Criteria
  • Agent Skills
Tutorials

Evaluation Tutorials

||View as Markdown|

Here are the hands-on walkthroughs for running benchmarks, collecting rollouts, and reading the outputs. They assume familiarity with the basic concepts in Evaluation and the workflow in Evaluation.

Evaluate EvalPlus

Run the EvalPlus coding benchmark and inspect rollout and aggregate metric outputs.

Reverify Rollouts

Recompute rewards from existing rollouts after changing a verifier parameter, without re-running model inference.

Workplace Assistant with Claude Code

Run an agentic tool-use benchmark end to end with the Claude Code agent harness — config, rollouts, and BLADE analysis.

notebook
Environment List

Browse the built-in benchmark and training environments.

Aggregate Metrics

Understand the aggregate metrics written after rollout collection.

Previous

Offline Training (SFT/DPO)

Next

Evaluate EvalPlus

NVIDIANVIDIA
Developer-friendly docs for your API
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2026, NVIDIA Corporation.