Browse Environments

View as Markdown

NeMo Gym includes 100+ environments covering math, coding, reasoning, knowledge, agentic tool use, instruction following, and safety. An environment includes a dataset, agent harness, verifier, and state. Some resources servers include agent composition and datasets and are runnable environments themselves; others are reusable scoring, tool, or state components and are not separate environment entries. Environments marked ✓ have a corresponding benchmark config with a canonical evaluation split and prepare_script.

Use the CLI to discover what’s available before running:

gym list benchmarks # all benchmarks with domain and description
gym list benchmarks aime24 # inspect one benchmark, including its agent harness
gym list environments # unified environment and benchmark catalog
gym search math # search environments and benchmarks by capability area
gym list benchmarks --json # machine-readable output for scripting or CI

Math & Science

EnvironmentDescriptionVerificationBenchmark
math_with_judgeOpenMathReasoning, DAPO, and MathStackOverflow datasetsmath-verify + LLM judge✓
math_with_autograderHard math benchmarks (e.g. IMO AnswerBench)math-verify + LLM autograder✓
polymathMultilingual math across 18 languages and 4 difficulty tiersLLM judge (weighted)✓
imo_gradingbenchFour-class IMO proof gradingLast-word extraction✓
imo_proofbench_judgeIMO ProofBench with 0–7 rubricLLM judge✓
math_proof_judgementBinary proof judgement — model reads a problem and proof, outputs Yes/NoLLM judge✓
math_formal_leanLean4 formal proof verificationLean4 compiler✓
physics_judgeOpen-ended physics QALLM judge + math-verify✓
ugphysics_judgeUndergraduate physics benchmarksLLM judge (TRUE/FALSE) + math-verify✓
critptResearch-level physics problems scored by the Artificial Analysis APILLM judge (external)✓
ether0Chemistry benchmark verifiers (ether0)Rule-based + LLM judge✓
bunsenbench_chemistry_mcqBunsenBench chemistry multiple-choice benchmarkExact match✓
assaybenchAssayBench: rank the hit genes of a CRISPR screen from its description (1,920 screens, 3 cohorts)Adjusted nDCG@100 via the authors’ assaybench package✓
frontierscience_judgeFrontierScience olympiad and research answer gradingLLM judge✓
proof_verificationProof scoring against ground truthLLM judge + meta-verifier
newton_benchScientific law discovery across 12 physics domainsExecution
math_with_codeCompetitive math with calculator toolsBoxed answer + numeric match

Coding

EnvironmentDescriptionVerificationBenchmark
bigcodebenchBigCodeBench Python solutions against unittest suiteCode execution✓
evalplusHumanEval+ and MBPP+ function completionCode execution✓
code_genCompetitive coding problem solvingCode execution✓
competitive_coding_challengesContest-style programming problemsCode execution✓
code_fimCode fill-in-the-middle (HumanEval-Infilling)Code execution✓
bird_sqlText-to-SQL on BIRD dev (1,534 SQLite tasks)SQL result-set equality✓
spider2_liteText-to-SQL on Spider 2.0-Lite (135 enterprise tasks)SQL result-set equality✓
text_to_sqlText-to-SQL across multiple SQL dialectsLLM judge (SQL equivalence)✓
swerl_genSWE patch and test generation in a sandboxed environmentCode execution (pytest)✓
scicodeMulti-step scientific code generationCode execution✓
cvdpRTL hardware design code generationCode execution (simulation)✓
openenvPython code execution environment via OpenEnvExecution (stdout/stderr)

Knowledge & Reasoning

EnvironmentDescriptionVerificationBenchmark
gpqa_diamondGraduate-level science multiple choice (GPQA Diamond)Exact match✓
mcqaMultiple-choice QA covering MMLU, GPQA, HLE, LongBenchExact match✓
equivalence_llm_judgeShort-answer QA with LLM-as-a-judge equivalence scoringLLM judge✓
hotpotqa_qaClosed-book multi-hop QA (HotPotQA)SQuAD-style substring match✓
simpleqaShort-form factual QA with abstention scoringLLM judge (3-tier)✓
omniscienceFactual recall and calibration QALLM judge✓
labbench2_vlmScientific VLM QA: figures, tables, lab protocolsLLM judge✓
litmus_agentLitmus-Bench v0.1 short-answer chemical reasoningAnswer extraction + numeric/boolean match✓
arc_agiAbstract reasoning puzzles (ARC-AGI)Exact match (grid)✓
nvarcARC-AGI in inductive (Python) and transductive (grid) modesCode execution / exact match✓
reasoning_gym100+ tasks: algebra, logic, geometry, graph theory, gamesExact match
multichallengeMulti-turn inference memory and instruction retentionLLM judge (rubric)
mrcrMulti-round coreference resolutionF1 (SequenceMatcher)✓
abstentionTrain models to abstain when unsure (HotPotQA, three-tier reward)LLM judge

Agentic / Tool Use

EnvironmentDescriptionVerificationBenchmark
workplace_assistantWorkplace tasks: 27 tools, 5 databases, 690 tasksRule-based (task completion)✓
aviaryMulti-hop QA with Wikipedia search + GSM8k with calculatorLLM judge + execution✓
tavily_searchWeb search tool use (Tavily API)Execution + optional LLM judge✓
calendarMulti-turn calendar scheduling with constraint satisfactionRule-based (constraints)✓
finance_sec_searchSEC EDGAR filing search for financial analysisLLM judge + execution✓
xlam_fcFunction calling from Salesforce xlam-60kExact match✓
single_step_tool_use_with_argument_comparisonPivot RL for tool use across conversational, SWE, and search domainsArgument comparison✓
talesText-adventure games (AlfWorld, ScienceWorld, Jericho, TextWorld) in Gymnasium API styleRule-based (task completion)
google_searchMCQA with integrated Google search toolExact match

Instruction Following & Safety

EnvironmentDescriptionVerificationBenchmark
ifbenchIFBench with 57 instruction types (AllenAI library)LLM judge✓
inverse_ifInverse instruction-following benchmark — instructions that counter conventional expectationsLLM judge (per-task)✓
structured_outputsSchema adherence across structured output formatsRule-based (schema validation)✓
indirect_prompt_injectionResistance to injected instructions in tool-use trajectoriesRule-based (attack detection)✓
jailbreak_detectionJailbreak resistance with Nemotron judgeLLM judge✓
xstestOver-refusal calibration (XSTest)Rule-based✓
instruction_followingIFEval and IFBench-style instruction followingLLM judge
verififVerifIF instruction-following validators (rule-based + LLM judge)Rule-based + LLM judge
format_verificationCitation format and freeform text formattingRegex / rule-based
structevalStructEval: JSON, YAML, CSV, TOML, XML schema adherenceRule-based (schema parsing)
over_refusal_detectionTrain models to avoid refusing safe promptsLLM judge

Other

EnvironmentDescriptionVerificationBenchmark
asr_with_pcASR with WER (standard, case-sensitive, punctuation)Execution (WER metrics)✓
wmt_translationMachine translation with BLEU and xCOMET-XXLExecution (COMET metrics)✓
longmt_evalDocument-level translation (SEGALE pipeline + COMETKiwi)Execution (neural QE)✓
vlm_eval_kitVLM benchmarks: MMBench, OCRBench, and othersExecution (VLMEvalKit)✓
graphwalksLong-context graph BFS/DFS reasoningF1 over node sets✓
arena_judgeArena Hard v2 pairwise LLM-judge (category-specific, side-swapped)LLM judge (pairwise)✓
speed_benchSpeculative-decoding throughput measurementvLLM Prometheus metrics✓
genrm_compareGenRM pairwise comparison for RLHF trainingLLM judge (pairwise)
blackjackGymnasium-style Blackjack (multi-step)Win/draw/loss
grl_sokobanSingle-box Sokoban puzzleExecution (puzzle solved)
grl_tetrisTetris in Gymnasium API styleRule-based (score/lines)