Appearance
Evaluation and Benchmarks
In a nutshell: this page covers evaluation and benchmarks in RL — evaluating RL is an order of magnitude harder than evaluating supervised learning: returns are random variables, results depend on random seeds, and every comparison drags in compute budgets. By the end you'll be able to set up a sound evaluation protocol ("fixed seeds × multiple runs × learning curves"), know what the mainstream benchmarks (Atari, MuJoCo, and friends) each measure, and spot cheats like "we ran one seed and called it a result".
1. Why RL evaluation is so hard
1.1 Supervised learning vs. RL: a fundamental difference in what's being evaluated
| Dimension | Supervised learning | RL |
|---|---|---|
| What is evaluated | A trained model | A policy interacting with a stochastic environment |
| Metric | A deterministic metric on a test set | Return (itself a random variable) |
| One evaluation | Run it once | Must run many times (both policy and environment are stochastic) |
| Environment | A fixed test set | Training environment = evaluation environment (no held-out test set) |
| Extra dimensions | None | Sample efficiency, compute budget |
The key point: in supervised learning, the train/test split is built in; in RL, the environment you train on and the environment you evaluate on are one and the same — the agent can simply "memorize" the state distribution of its training environment and still score well at evaluation time (this is one source of "evaluation cheating"; see Section 4).
1.2 Four sources of variance in returns
A single RL evaluation yields one return number, polluted by four layers of randomness at once:
| Randomness source | Explanation |
|---|---|
| Environment stochasticity | The same policy follows different trajectories even from identical initial conditions |
| Policy stochasticity | Stochastic policies (SAC/PPO) sample different actions each time |
| The seed chain | The initial seed determines network initialization + data ordering + the exploration sequence |
| Training noise | Mini-batch sampling and gradient noise make every training run different |
So the return of "one run" proves nothing. The correct unit is the distribution over multiple runs (multiple seeds).
1.3 Non-stationarity: evaluating during training vs. after training
An RL policy keeps changing throughout training, so evaluation must specify which point in time is being evaluated:
- Final performance: the level of the policy at the end of training;
- Learning curve: how performance evolves over the whole course of training (the single most important visualization for tuning; see Hyperparameter Tuning and Optimization).
2. The elements of an evaluation protocol: five must-haves
Any defensible RL evaluation includes at least the following:
text
Five elements of RL evaluation
─────────────────────────────
1. A fixed set of seeds (at least 5-10, spanning 0..n)
2. Full training for every seed (no cherry-picking "the best checkpoint"
along the way)
3. Multiple evaluation episodes averaged at each evaluation point
(to average out policy/environment noise)
4. Report a distribution: median + IQR, not mean ± std
5. State the budget explicitly: training steps / environment
interactions / compute resources
─────────────────────────────2.1 Why median + IQR instead of mean ± standard deviation
Return distributions in RL are usually long-tailed (an occasional explosive high score, or an occasional collapse to near zero), and the mean gets dragged around by the tails. The median is robust, and the IQR (interquartile range) shows stability at a glance. Example reporting format:
Algorithm Median return IQR Median steps to success Interactions
PPO 245.0 [221, 262] 1.2M 50M
SAC 398.5 [372, 405] 0.8M 50M2.2 Learning curves: the core visualization of evaluation
X-axis = training steps (or environment interactions), Y-axis = return. Three typical shapes:
text
Not converging (undertrained) Oscillating (unstable) Normal convergence
▲ ▲ ▲
│ ╱╲ │ ╱╲╱╲╱ │ ╱‾‾‾
│╱ ╲ │╱ ╲ │ ╱
│ ╲ │ ╲╱╲ │ ╱
│ ╲ │ ╲ │╱
└─────────▶ └─────────▶ └─────────▶- Not converging: a reward-design or hyperparameter problem (check Reward Engineering before blaming the hyperparameters);
- Oscillating: learning rate too high / too much exploration / value overestimation (see The Actor–Critic Family);
- Normal: flat at first → steep rise in the middle → plateau at the end.
3. Sample efficiency vs. final performance: two dimensions you must look at separately
3.1 A common illusion
"Algorithm A's final return is higher than B's, so A is better" — wrong. A may have used 10× the samples to get there. Where samples are expensive (robotics, real business systems), sample efficiency matters just as much as final performance.
| Dimension | Metric | Typical comparison |
|---|---|---|
| Sample efficiency | Interactions needed to reach a performance threshold | SAC reaches score X in 100K steps; PPO needs 1M steps |
| Final performance | Best return after convergence | PPO may end up with the higher final score |
A proper evaluation report plots both axes at once: interactions on the x-axis, so you can see which curve climbs first and which climbs highest. See Building an RL Evaluation from Scratch.
3.2 Fairness of compute budgets
On-policy (PPO) and off-policy (SAC) methods have different compute costs for the same number of "environment interactions" (PPO must run both the actor and critic networks on every rollout sample). A report must declare all three: environment interactions + wall-clock training time + GPU/CPU resources. Reporting only one of them is "unfair budgeting".
Common forms of evaluation cheating
- Reporting "the best checkpoint during training" (not the final policy);
- Running a single seed that happens to be a lucky one;
- Tuning your own algorithm's hyperparameters while leaving the competitor at defaults ("apples to oranges");
- Evaluating on the training environment (memorized answers). The full anti-pattern list is in Common Pitfalls and Anti-Patterns.
4. A map of mainstream benchmarks
4.1 Benchmarks at a glance
| Benchmark | Environments / tasks | Action space | Difficulty | What it measures | Best for |
|---|---|---|---|---|---|
| Gymnasium | Classic control (CartPole, MountainCar, LunarLander) | Discrete / continuous | Low | Getting started, debugging | Learning, algorithm sanity checks |
| Atari 100k / ALE | 49+ arcade games (pixel input) | Discrete (18) | Medium–high | Vision, sparse rewards, generalization | The deep-RL standard |
| MuJoCo | Robotic joint control (HalfCheetah, Humanoid) | Continuous | Medium | Continuous control | The default for algorithm comparisons |
| DM Control | Pixel/state continuous control | Continuous | Medium–high | Vision-based control | The DeepMind lineage |
| Procgen | 16 procedurally generated games | Discrete | Medium | Generalization (unseen levels) | Generalization research |
| Meta-World | 50 robotic manipulation tasks | Continuous | Medium–high | Multi-task, generalization | Meta-RL |
| MiniGrid / Maze | Grid-world mazes | Discrete | Low–medium | Exploration, sparse rewards | Exploration research |
| D4RL | Offline datasets | Discrete / continuous | Medium | Offline RL | Offline algorithms |
| Isaac Gym / Brax | Massively parallel physics | Continuous | High | Large-scale parallelism | Large-scale training |
4.2 Two pillars: Atari and MuJoCo
- Atari (Arcade Learning Environment, Bellemare et al., 2013): learn policies from raw pixels — a test of "vision + exploration + temporal reasoning", and the main battlefield of the DQN family and exploration research. Criticisms: some games score high under a random policy; the environments are highly deterministic (easy to memorize); score scales vary wildly (human-normalized scores are needed).
- MuJoCo (Todorov et al., 2012): the continuous-control benchmark — physics simulation, precisely reproducible, and the standard arena for SAC/TD3/PPO. Criticisms: a narrow task set; easy to "overfit the benchmark" (tuning that only works on these tasks); dynamics far from reality.
4.3 Generalization benchmarks: Procgen and Meta-World
Supervised learning cares about generalization; RL used not to. Generalization benchmarks correct this:
- Procgen: each training seed generates different levels, and the training and test levels do not overlap — a direct test of "can the policy transfer to unseen levels";
- Meta-World: 50 distinct robotic-arm manipulation tasks, testing "multi-task sharing and fast adaptation".
Engineering advice on choosing benchmarks
- Getting started / teaching: Gymnasium (see Progressive Tutorials);
- Validating a continuous-control algorithm: MuJoCo, starting with 3 seeds;
- Validating a discrete/vision algorithm: Atari 100k (few samples, fast iteration);
- Claiming "it generalizes": Procgen / Meta-World;
- Offline algorithms: D4RL (for the corresponding datasets, see Datasets and Tools).
5. Evaluation reporting standards: what a credible RL experiment should report
5.1 Reporting checklist
text
What a complete RL evaluation report should include
────────────────────────────
Environment & task: version numbers, reward design,
truncation/termination definitions
Number of seeds: at least 5 (10+ for papers)
Evaluation protocol: training steps, evaluation interval,
episodes per evaluation
Metrics: median + IQR + learning curves (curves for ALL seeds)
Budget: interactions, wall-clock time, hardware
Hyperparameters: full table + search method (random / grid / Bayesian)
Code: a pointer to a reproducible implementation (or an appendix)
────────────────────────────5.2 Visualization standards for learning curves
- Plot the median as the curve and the IQR as the shaded band (don't plot the mean alone);
- Use "environment interactions" as the unified x-axis (different algorithms count steps differently — a PPO step is not a SAC step);
- Include a "training time vs. performance" comparison (compute fairness);
- Plot all algorithms on the same chart with the same scales.
How this connects to the practice pages
This page covers "why evaluation is hard and what benchmarks measure"; for building an evaluation protocol hands-on and report templates, see Building an RL Evaluation from Scratch; for using evaluation to separate "hyperparameter problems vs. algorithm problems" during tuning, see Hyperparameter Tuning and Optimization.
6. Critiques of benchmarks and the "benchmark overfitting" problem
6.1 Benchmark overfitting: "teaching to the test" in RL
"SOTA on HalfCheetah" says less and less about real capability. Why:
- Too few tasks: mainstream comparisons concentrate on about a dozen tasks, and hyperparameters can be tuned to "favor exactly this dozen" — swap in a different set of tasks and the mirage vanishes;
- Evaluation is training: the training and evaluation environments are the same, so the agent can memorize the state distribution (see Section 1);
- Publication incentives: papers are judged by "+N points on benchmark X", so research chases benchmarks instead of problems.
This is benchmark overfitting: the score goes up, generalization doesn't.
6.2 Practices that reduce overfitting
| Technique | Explanation |
|---|---|
| Prefer procedurally generated environments | Procgen makes every level different, with non-overlapping train/test levels (see Datasets and Tools) |
| Cross-validate across environments | Don't report MuJoCo alone — mix Atari + DM Control + your own tasks |
| Report the "tuning cost" | State clearly whether the hyperparameters are defaults or were specifically searched for (honest disclosure of fairness) |
| A held-out test distribution | Deliberately reserve a set of tasks that "never participated in development" for final evaluation |
6.3 The right mindset toward benchmarks
A benchmark is not a tool to "prove your algorithm is good" — it is an instrument for "exposing your algorithm's weaknesses". Scores are for reviewers; diagnostics are for you — spending the same hours staring at "why does this curve stop improving after 1M steps" is worth more than running yet another benchmark.
7. The exploration dimension of evaluation: don't forget "in-training behavior"
Evaluation is not just "the final score". There is another critical dimension during training — whether exploration stays healthy (see Exploration and Exploitation). When tuning, plot three curves at once:
| Curve | What to look for |
|---|---|
| Return | Whether performance is rising |
| Policy entropy | Whether exploration is being squeezed shut too early (an entropy cliff = premature convergence) |
| Q-values | Overestimation / inflation (Q climbs while return doesn't = value hallucination) |
These three curves are the "diagnostic dashboard" of Hyperparameter Tuning and Optimization.
8. Evaluation emphases across subfields
Evaluation is not one-size-fits-all. Offline RL, multi-agent, and RLHF each have their own "evaluation is hard" — an evaluation that ignores these specifics is worse than no evaluation.
8.1 Offline RL: the hardest class to evaluate
Offline RL forbids interaction during training, so "is this new policy any good?" can only be answered offline. The available tools, and their pitfalls:
| Technique | How it works | Main problem |
|---|---|---|
| Trajectory-wise importance sampling | Estimate by reweighting in-data trajectories with IS | Variance explodes exponentially with trajectory length — nearly unusable for long-horizon tasks |
| Scoring with the learned Q/critic | Evaluate the policy with the trained value function | Q itself hallucinates on OOD actions (see Offline RL) — "judging hallucinations with hallucinations" |
| World-model simulation | Learn an environment model and evaluate inside it | Model error propagates into the evaluation conclusion |
| Conservative values (the CQL family) | Rank policies with values that "never overestimate" | Gives only a relative lower bound, not an exact value |
Engineering conclusion: the only reliable use of offline evaluation is "ranking candidate policies", followed by a small-scale online validation. Treat any scheme claiming to "predict absolute online performance offline" with suspicion (see Common Pitfalls and Anti-Patterns).
8.2 Multi-agent: the opponent is the "test set"
The biggest problem in multi-agent evaluation: the result depends on who the opponent is. A policy that beats opponent A may lose to opponent B. Evaluation must answer three questions:
- Against whom: fixed baseline opponents (random/fixed policies), best-response (the strongest opponent), or self-play (playing against yourself)?
- Coalition stability: after co-training with one partner, does the policy still work with a different partner?
- Emergent capabilities: do human-interpretable high-level behaviors appear? (Both OpenAI Five and AlphaStar were evaluated through a dual channel of "win rate + human review".)
A common trap: reporting only "win rate against a random policy" — that only proves the agent isn't acting randomly; it proves nothing about real strength. A rigorous MARL paper reports the "full win-rate matrix against multiple baselines"; see Multi-Agent RL.
8.3 RLHF / language models: three metric families, each with its own agenda
Evaluating LLM alignment means watching three dimensions at once; any single metric will be gamed:
| Metric | What it measures | Known weakness |
|---|---|---|
| Reward model score | Alignment with a proxy for human preference | Goodhart's law: optimized to the extreme, it drifts from real preferences (see RLHF and Alignment with Human Feedback) |
| Human evaluation (preference win rate) | Real human preferences | Expensive, subjective, hard to reproduce |
| Standard capability benchmarks (MMLU, code evals) | Whether general capability has dropped | Measures "capability", not "alignment" |
Engineering rule of thumb: reward score up while human win rate is flat or falling = overoptimization in progress; stop immediately. This is also how the alignment tax is monitored in LLM Alignment in Practice: RLHF.
8.4 Quantifying learning curves: don't rely on "eyeballing" alone
"Eyeballing which curve is higher" is unreliable. Two standard practices:
- AUC (area under the curve): integrate the learning curve, penalizing both "slow" and "low" — one metric that captures sample efficiency and final performance together;
- IQM (interquartile mean): average the middle 50% of returns across runs — more robust than the mean and more informative than the median (recommended by Agarwal et al., 2021).
9. A checklist of common mistakes
- Drawing conclusions from 1 seed — that's a luck curve, nothing more;
- Reporting the best checkpoint — not the true level of the policy;
- Leaving the competitor at default hyperparameters — defaults are not fairness;
- Not reporting the compute budget — sample efficiency becomes unjudgeable;
- Evaluating on the training environment — memorized-answer scores;
- Treating the mean as everything — it gets dragged around by long tails;
- Confusing "steps" across algorithms — on-policy and off-policy count steps differently; everything must be converted to "environment interactions".
Frequent interview question: "How do you evaluate your RL project?"
Answer: fix at least 5 seeds, train each to completion, report the median and IQR of final returns, plot learning curves (x-axis: environment interactions), and report wall-clock time plus a full hyperparameter table. Then turn the question back on the interviewer: "What is the evaluation budget and goal?" — the evaluation protocol itself should fit the project's goal (shipping decisions need small-sample fast evaluation; papers need large-scale rigorous evaluation).
Further Reading
- Building an RL Evaluation from Scratch — building an evaluation protocol hands-on, with report templates
- Datasets and Tools — an index of tools and datasets: Gymnasium, MuJoCo, Atari, D4RL, and more
- Exploration and Exploitation — evaluating exploration health during training
- Common Pitfalls and Anti-Patterns — detecting evaluation cheating and seed overfitting
- Hyperparameter Tuning and Optimization — learning-curve shape analysis and multi-seed confirmation
References
- Henderson, P., Islam, R., Bachman, P., et al. (2018). Deep Reinforcement Learning that Matters. AAAI. arXiv:1709.06560 (the classic empirical paper on the unreliability of RL evaluation — a must-read)
- Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The Arcade Learning Environment: An Evaluation Platform for General Agents. JAIR. The original Atari benchmark paper.
- Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A physics engine for model-based control. IROS. The MuJoCo engine paper.
- Cobbe, K., Hesse, C., Hilton, J., & Schulman, J. (2020). Leveraging Procedural Generation to Benchmark Reinforcement Learning. ICML. arXiv:1912.01588 (Procgen)
- Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2020). D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv:2004.07219 (the D4RL offline benchmark)
- Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., & Bellemare, M. G. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS. arXiv:2108.13264 (reliable RL evaluation with aggregate statistics — the authoritative source for IQR reporting)
- Gymnasium official docs: https://gymnasium.farama.org/