Evaluation Tutorials
Here are the hands-on walkthroughs for running benchmarks, collecting rollouts, and reading the outputs. They assume familiarity with the basic concepts in Evaluation and the workflow in Evaluation.
Evaluate EvalPlus
Run the EvalPlus coding benchmark and inspect rollout and aggregate metric outputs.
Reverify Rollouts
Recompute rewards from existing rollouts after changing a verifier parameter, without re-running model inference.
Workplace Assistant with Claude Code
Run an agentic tool-use benchmark end to end with the Claude Code agent harness — config, rollouts, and BLADE analysis.
notebookEnvironment List
Browse the built-in benchmark and training environments.
Aggregate Metrics
Understand the aggregate metrics written after rollout collection.