Skip to content

Evaluation and Benchmarks

On this page Why RL evaluation is hard — return variance, random seeds, sample efficiency, and compute budgets; a map of the mainstream benchmarks (Gymnasium, Atari, MuJoCo, DM Control, Procgen, Meta-World, Maze); evaluation reporting standards.

Evaluation and Benchmarks ​

In a nutshell: this page covers evaluation and benchmarks in RL — evaluating RL is an order of magnitude harder than evaluating supervised learning: returns are random variables, results depend on random seeds, and every comparison drags in compute budgets. By the end you'll be able to set up a sound evaluation protocol ("fixed seeds × multiple runs × learning curves"), know what the mainstream benchmarks (Atari, MuJoCo, and friends) each measure, and spot cheats like "we ran one seed and called it a result".

1. Why RL evaluation is so hard ​

1.1 Supervised learning vs. RL: a fundamental difference in what's being evaluated ​

DimensionSupervised learningRL
What is evaluatedA trained modelA policy interacting with a stochastic environment
MetricA deterministic metric on a test setReturn (itself a random variable)
One evaluationRun it onceMust run many times (both policy and environment are stochastic)
EnvironmentA fixed test setTraining environment = evaluation environment (no held-out test set)
Extra dimensionsNoneSample efficiency, compute budget

The key point: in supervised learning, the train/test split is built in; in RL, the environment you train on and the environment you evaluate on are one and the same — the agent can simply "memorize" the state distribution of its training environment and still score well at evaluation time (this is one source of "evaluation cheating"; see Section 4).

1.2 Four sources of variance in returns ​

A single RL evaluation yields one return number, polluted by four layers of randomness at once:

Randomness sourceExplanation
Environment stochasticityThe same policy follows different trajectories even from identical initial conditions
Policy stochasticityStochastic policies (SAC/PPO) sample different actions each time
The seed chainThe initial seed determines network initialization + data ordering + the exploration sequence
Training noiseMini-batch sampling and gradient noise make every training run different

So the return of "one run" proves nothing. The correct unit is the distribution over multiple runs (multiple seeds).

1.3 Non-stationarity: evaluating during training vs. after training ​

An RL policy keeps changing throughout training, so evaluation must specify which point in time is being evaluated:

  • Final performance: the level of the policy at the end of training;
  • Learning curve: how performance evolves over the whole course of training (the single most important visualization for tuning; see Hyperparameter Tuning and Optimization).

2. The elements of an evaluation protocol: five must-haves ​

Any defensible RL evaluation includes at least the following:

text
Five elements of RL evaluation
─────────────────────────────
1. A fixed set of seeds (at least 5-10, spanning 0..n)
2. Full training for every seed (no cherry-picking "the best checkpoint"
   along the way)
3. Multiple evaluation episodes averaged at each evaluation point
   (to average out policy/environment noise)
4. Report a distribution: median + IQR, not mean ± std
5. State the budget explicitly: training steps / environment
   interactions / compute resources
─────────────────────────────

2.1 Why median + IQR instead of mean ± standard deviation ​

Return distributions in RL are usually long-tailed (an occasional explosive high score, or an occasional collapse to near zero), and the mean gets dragged around by the tails. The median is robust, and the IQR (interquartile range) shows stability at a glance. Example reporting format:

Algorithm   Median return   IQR          Median steps to success   Interactions
PPO         245.0           [221, 262]   1.2M                      50M
SAC         398.5           [372, 405]   0.8M                      50M

2.2 Learning curves: the core visualization of evaluation ​

X-axis = training steps (or environment interactions), Y-axis = return. Three typical shapes:

text
Not converging (undertrained)   Oscillating (unstable)    Normal convergence
▲                               ▲                         ▲
│ ╱╲                           │ ╱╲╱╲╱                  │    ╱‾‾‾
│╱  ╲                          │╱      ╲                 │  ╱
│    ╲                         │         ╲╱╲             │ ╱
│     ╲                        │              ╲          │╱
└─────────▶                    └─────────▶                └─────────▶
  • Not converging: a reward-design or hyperparameter problem (check Reward Engineering before blaming the hyperparameters);
  • Oscillating: learning rate too high / too much exploration / value overestimation (see The Actor–Critic Family);
  • Normal: flat at first → steep rise in the middle → plateau at the end.

3. Sample efficiency vs. final performance: two dimensions you must look at separately ​

3.1 A common illusion ​

"Algorithm A's final return is higher than B's, so A is better" — wrong. A may have used 10× the samples to get there. Where samples are expensive (robotics, real business systems), sample efficiency matters just as much as final performance.

DimensionMetricTypical comparison
Sample efficiencyInteractions needed to reach a performance thresholdSAC reaches score X in 100K steps; PPO needs 1M steps
Final performanceBest return after convergencePPO may end up with the higher final score

A proper evaluation report plots both axes at once: interactions on the x-axis, so you can see which curve climbs first and which climbs highest. See Building an RL Evaluation from Scratch.

3.2 Fairness of compute budgets ​

On-policy (PPO) and off-policy (SAC) methods have different compute costs for the same number of "environment interactions" (PPO must run both the actor and critic networks on every rollout sample). A report must declare all three: environment interactions + wall-clock training time + GPU/CPU resources. Reporting only one of them is "unfair budgeting".

Common forms of evaluation cheating

  1. Reporting "the best checkpoint during training" (not the final policy);
  2. Running a single seed that happens to be a lucky one;
  3. Tuning your own algorithm's hyperparameters while leaving the competitor at defaults ("apples to oranges");
  4. Evaluating on the training environment (memorized answers). The full anti-pattern list is in Common Pitfalls and Anti-Patterns.

4. A map of mainstream benchmarks ​

4.1 Benchmarks at a glance ​

BenchmarkEnvironments / tasksAction spaceDifficultyWhat it measuresBest for
GymnasiumClassic control (CartPole, MountainCar, LunarLander)Discrete / continuousLowGetting started, debuggingLearning, algorithm sanity checks
Atari 100k / ALE49+ arcade games (pixel input)Discrete (18)Medium–highVision, sparse rewards, generalizationThe deep-RL standard
MuJoCoRobotic joint control (HalfCheetah, Humanoid)ContinuousMediumContinuous controlThe default for algorithm comparisons
DM ControlPixel/state continuous controlContinuousMedium–highVision-based controlThe DeepMind lineage
Procgen16 procedurally generated gamesDiscreteMediumGeneralization (unseen levels)Generalization research
Meta-World50 robotic manipulation tasksContinuousMedium–highMulti-task, generalizationMeta-RL
MiniGrid / MazeGrid-world mazesDiscreteLow–mediumExploration, sparse rewardsExploration research
D4RLOffline datasetsDiscrete / continuousMediumOffline RLOffline algorithms
Isaac Gym / BraxMassively parallel physicsContinuousHighLarge-scale parallelismLarge-scale training

4.2 Two pillars: Atari and MuJoCo ​

  • Atari (Arcade Learning Environment, Bellemare et al., 2013): learn policies from raw pixels — a test of "vision + exploration + temporal reasoning", and the main battlefield of the DQN family and exploration research. Criticisms: some games score high under a random policy; the environments are highly deterministic (easy to memorize); score scales vary wildly (human-normalized scores are needed).
  • MuJoCo (Todorov et al., 2012): the continuous-control benchmark — physics simulation, precisely reproducible, and the standard arena for SAC/TD3/PPO. Criticisms: a narrow task set; easy to "overfit the benchmark" (tuning that only works on these tasks); dynamics far from reality.

4.3 Generalization benchmarks: Procgen and Meta-World ​

Supervised learning cares about generalization; RL used not to. Generalization benchmarks correct this:

  • Procgen: each training seed generates different levels, and the training and test levels do not overlap — a direct test of "can the policy transfer to unseen levels";
  • Meta-World: 50 distinct robotic-arm manipulation tasks, testing "multi-task sharing and fast adaptation".

Engineering advice on choosing benchmarks

  • Getting started / teaching: Gymnasium (see Progressive Tutorials);
  • Validating a continuous-control algorithm: MuJoCo, starting with 3 seeds;
  • Validating a discrete/vision algorithm: Atari 100k (few samples, fast iteration);
  • Claiming "it generalizes": Procgen / Meta-World;
  • Offline algorithms: D4RL (for the corresponding datasets, see Datasets and Tools).

5. Evaluation reporting standards: what a credible RL experiment should report ​

5.1 Reporting checklist ​

text
What a complete RL evaluation report should include
────────────────────────────
Environment & task: version numbers, reward design,
  truncation/termination definitions
Number of seeds: at least 5 (10+ for papers)
Evaluation protocol: training steps, evaluation interval,
  episodes per evaluation
Metrics: median + IQR + learning curves (curves for ALL seeds)
Budget: interactions, wall-clock time, hardware
Hyperparameters: full table + search method (random / grid / Bayesian)
Code: a pointer to a reproducible implementation (or an appendix)
────────────────────────────

5.2 Visualization standards for learning curves ​

  • Plot the median as the curve and the IQR as the shaded band (don't plot the mean alone);
  • Use "environment interactions" as the unified x-axis (different algorithms count steps differently — a PPO step is not a SAC step);
  • Include a "training time vs. performance" comparison (compute fairness);
  • Plot all algorithms on the same chart with the same scales.

How this connects to the practice pages

This page covers "why evaluation is hard and what benchmarks measure"; for building an evaluation protocol hands-on and report templates, see Building an RL Evaluation from Scratch; for using evaluation to separate "hyperparameter problems vs. algorithm problems" during tuning, see Hyperparameter Tuning and Optimization.

6. Critiques of benchmarks and the "benchmark overfitting" problem ​

6.1 Benchmark overfitting: "teaching to the test" in RL ​

"SOTA on HalfCheetah" says less and less about real capability. Why:

  • Too few tasks: mainstream comparisons concentrate on about a dozen tasks, and hyperparameters can be tuned to "favor exactly this dozen" — swap in a different set of tasks and the mirage vanishes;
  • Evaluation is training: the training and evaluation environments are the same, so the agent can memorize the state distribution (see Section 1);
  • Publication incentives: papers are judged by "+N points on benchmark X", so research chases benchmarks instead of problems.

This is benchmark overfitting: the score goes up, generalization doesn't.

6.2 Practices that reduce overfitting ​

TechniqueExplanation
Prefer procedurally generated environmentsProcgen makes every level different, with non-overlapping train/test levels (see Datasets and Tools)
Cross-validate across environmentsDon't report MuJoCo alone — mix Atari + DM Control + your own tasks
Report the "tuning cost"State clearly whether the hyperparameters are defaults or were specifically searched for (honest disclosure of fairness)
A held-out test distributionDeliberately reserve a set of tasks that "never participated in development" for final evaluation

6.3 The right mindset toward benchmarks ​

A benchmark is not a tool to "prove your algorithm is good" — it is an instrument for "exposing your algorithm's weaknesses". Scores are for reviewers; diagnostics are for you — spending the same hours staring at "why does this curve stop improving after 1M steps" is worth more than running yet another benchmark.

7. The exploration dimension of evaluation: don't forget "in-training behavior" ​

Evaluation is not just "the final score". There is another critical dimension during training — whether exploration stays healthy (see Exploration and Exploitation). When tuning, plot three curves at once:

CurveWhat to look for
ReturnWhether performance is rising
Policy entropyWhether exploration is being squeezed shut too early (an entropy cliff = premature convergence)
Q-valuesOverestimation / inflation (Q climbs while return doesn't = value hallucination)

These three curves are the "diagnostic dashboard" of Hyperparameter Tuning and Optimization.

8. Evaluation emphases across subfields ​

Evaluation is not one-size-fits-all. Offline RL, multi-agent, and RLHF each have their own "evaluation is hard" — an evaluation that ignores these specifics is worse than no evaluation.

8.1 Offline RL: the hardest class to evaluate ​

Offline RL forbids interaction during training, so "is this new policy any good?" can only be answered offline. The available tools, and their pitfalls:

TechniqueHow it worksMain problem
Trajectory-wise importance samplingEstimate by reweighting in-data trajectories with ISVariance explodes exponentially with trajectory length — nearly unusable for long-horizon tasks
Scoring with the learned Q/criticEvaluate the policy with the trained value functionQ itself hallucinates on OOD actions (see Offline RL) — "judging hallucinations with hallucinations"
World-model simulationLearn an environment model and evaluate inside itModel error propagates into the evaluation conclusion
Conservative values (the CQL family)Rank policies with values that "never overestimate"Gives only a relative lower bound, not an exact value

Engineering conclusion: the only reliable use of offline evaluation is "ranking candidate policies", followed by a small-scale online validation. Treat any scheme claiming to "predict absolute online performance offline" with suspicion (see Common Pitfalls and Anti-Patterns).

8.2 Multi-agent: the opponent is the "test set" ​

The biggest problem in multi-agent evaluation: the result depends on who the opponent is. A policy that beats opponent A may lose to opponent B. Evaluation must answer three questions:

  • Against whom: fixed baseline opponents (random/fixed policies), best-response (the strongest opponent), or self-play (playing against yourself)?
  • Coalition stability: after co-training with one partner, does the policy still work with a different partner?
  • Emergent capabilities: do human-interpretable high-level behaviors appear? (Both OpenAI Five and AlphaStar were evaluated through a dual channel of "win rate + human review".)

A common trap: reporting only "win rate against a random policy" — that only proves the agent isn't acting randomly; it proves nothing about real strength. A rigorous MARL paper reports the "full win-rate matrix against multiple baselines"; see Multi-Agent RL.

8.3 RLHF / language models: three metric families, each with its own agenda ​

Evaluating LLM alignment means watching three dimensions at once; any single metric will be gamed:

MetricWhat it measuresKnown weakness
Reward model scoreAlignment with a proxy for human preferenceGoodhart's law: optimized to the extreme, it drifts from real preferences (see RLHF and Alignment with Human Feedback)
Human evaluation (preference win rate)Real human preferencesExpensive, subjective, hard to reproduce
Standard capability benchmarks (MMLU, code evals)Whether general capability has droppedMeasures "capability", not "alignment"

Engineering rule of thumb: reward score up while human win rate is flat or falling = overoptimization in progress; stop immediately. This is also how the alignment tax is monitored in LLM Alignment in Practice: RLHF.

8.4 Quantifying learning curves: don't rely on "eyeballing" alone ​

"Eyeballing which curve is higher" is unreliable. Two standard practices:

  • AUC (area under the curve): integrate the learning curve, penalizing both "slow" and "low" — one metric that captures sample efficiency and final performance together;
  • IQM (interquartile mean): average the middle 50% of returns across runs — more robust than the mean and more informative than the median (recommended by Agarwal et al., 2021).

9. A checklist of common mistakes ​

  1. Drawing conclusions from 1 seed — that's a luck curve, nothing more;
  2. Reporting the best checkpoint — not the true level of the policy;
  3. Leaving the competitor at default hyperparameters — defaults are not fairness;
  4. Not reporting the compute budget — sample efficiency becomes unjudgeable;
  5. Evaluating on the training environment — memorized-answer scores;
  6. Treating the mean as everything — it gets dragged around by long tails;
  7. Confusing "steps" across algorithms — on-policy and off-policy count steps differently; everything must be converted to "environment interactions".

Frequent interview question: "How do you evaluate your RL project?"

Answer: fix at least 5 seeds, train each to completion, report the median and IQR of final returns, plot learning curves (x-axis: environment interactions), and report wall-clock time plus a full hyperparameter table. Then turn the question back on the interviewer: "What is the evaluation budget and goal?" — the evaluation protocol itself should fit the project's goal (shipping decisions need small-sample fast evaluation; papers need large-scale rigorous evaluation).

Further Reading ​

References ​

  • Henderson, P., Islam, R., Bachman, P., et al. (2018). Deep Reinforcement Learning that Matters. AAAI. arXiv:1709.06560 (the classic empirical paper on the unreliability of RL evaluation — a must-read)
  • Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The Arcade Learning Environment: An Evaluation Platform for General Agents. JAIR. The original Atari benchmark paper.
  • Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A physics engine for model-based control. IROS. The MuJoCo engine paper.
  • Cobbe, K., Hesse, C., Hilton, J., & Schulman, J. (2020). Leveraging Procedural Generation to Benchmark Reinforcement Learning. ICML. arXiv:1912.01588 (Procgen)
  • Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2020). D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv:2004.07219 (the D4RL offline benchmark)
  • Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., & Bellemare, M. G. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS. arXiv:2108.13264 (reliable RL evaluation with aggregate statistics — the authoritative source for IQR reporting)
  • Gymnasium official docs: https://gymnasium.farama.org/