Appearance
Building an RL Evaluation Suite from Scratch
One-liner: this page teaches you to build an RL evaluation protocol that won't lie to you — from experiment-matrix design and metric selection to fairness pitfalls and the final report, all in one pass. It resolves the recurring puzzle of "training curves keep climbing, yet everything falls apart the moment real comparisons begin," and it serves anyone who has to make decisions from experimental results (tuning hyperparameters, choosing algorithms, deciding whether to ship). By the end, you'll have a reusable evaluation script and a report template of your own.
First, a hard truth: a large share of published RL results cannot be reproduced with a single run. Henderson et al., in "Deep Reinforcement Learning that Matters," put numbers to it: for the same algorithm from the same paper, performance variation across seeds can exceed the variation between different algorithms. Put differently, if you can't evaluate, your experimental conclusions are dice rolls. This page turns dice rolls into readings.
This page is the hands-on companion to Evaluation and Benchmarks: that concept page explains why RL evaluation is hard; this one shows you exactly how to build the evaluation itself.
1. Clarify First: What Question Must This Evaluation Answer
An evaluation protocol isn't better for having more pieces — it's better for serving a decision. Start by distinguishing three evaluation goals, because their protocol designs differ completely:
| Evaluation goal | The question it must answer | Key requirement | Typical scope |
|---|---|---|---|
| Hyperparameter decision | Does changing this hyperparameter help or hurt? | Relative comparison is enough; few seeds OK | Mostly single runs |
| Algorithm comparison | Is algorithm A really better than B? | Multiple seeds + statistical intervals; fairness is unforgiving | A multi-seed matrix |
| Launch decision | Can we ship? What's the cost of failure? | Tail-risk analysis, safety boundaries, offline + online validation | The heaviest; long cycles |
The most common evaluation mistake: applying a tuning-grade protocol to a launch decision
A "+5 points" bump during tuning may be nothing more than seed luck. Apply the same lax standard to a launch decision, and you've mistaken one draw of variance for a real signal. The more consequential the conclusion, the stricter the evaluation protocol. That's the principle from which every detail on this page follows.
2. Designing the Experiment Matrix: Algorithm x Seed x Environment
The basic unit of evaluation is a single run. A run is a quadruple (algorithm, seed, environment, config), and the full set of runs forms the matrix:
text
Experiment matrix (columns = seeds, rows = algorithm x environment)
seed 0 seed 1 seed 2 seed 3 seed 4
PPO CartPole ██████ ██████ ██████ ██████ ██████
PPO LunarLander ██████ ██████ ██████ ██████ ██████
SAC CartPole ██████ ██████ ██████ ██████ ██████
SAC LunarLander ██████ ██████ ██████ ██████ ██████
█ = one full training run + evaluation (learning curve included)1. How to Choose Seeds
| Element | Recommended practice | Why |
|---|---|---|
| Number of seeds | >= 5 for comparisons, >= 10 for formal reports | It takes about 5 seeds for the mean to settle to a distinguishable signal |
| Seed set | Fix one set: 0, 1, 2, 3, 4 | Sharing the same seed set across algorithms maximizes comparability |
| Seed values | Record the full config alongside each seed | Different seeds and different hyperparameters = uncontrolled variables, no attribution |
| Environment randomness | The seed must be threaded into the environment's reset and action_space | Otherwise the environment is only partially controlled — see Common Pitfalls |
2. The Variable-Control Checklist
When comparing two algorithms, "algorithm" is the only variable allowed to differ. Everything else must be locked down:
- [ ] Environment version and wrappers (is
TimeLimiton?frame_stack? identical across runs) - [ ] Per-seed environment randomness (each seed must seed the environment, not just one global seed at startup)
- [ ] Observation preprocessing (normalization, grayscale, frame stacking)
- [ ] Total training-step budget (not episodes or iterations — what an "epoch" means varies by algorithm, so the budget must be counted in environment interactions)
- [ ] The evaluation protocol itself (see Section 3)
- [ ] Software versions (pin the exact gymnasium/numpy/torch versions in the report)
Why the training budget must be counted in environment steps, not training iterations
One PPO update consumes 2,048 steps; one REINFORCE update consumes a whole episode (tens to hundreds of steps). Align budgets by "number of updates" and PPO gets a massive, invisible advantage. Align them by environment steps, and both sides spend the same money to buy the same amount of data — a fair comparison. This is also the premise for sample efficiency vs. final performance.
3. Metrics: Mean Return, Learning Curves, Success Rate
1. The Learning Curve Is the Mother of All Metrics
Don't just report a "final score." The learning curve (x-axis: environment steps, y-axis: evaluation return) preserves the time dimension: who converges fast, who's degrading, who's oscillating early on. A single learning-curve plot with an IQR band across seeds carries an order of magnitude more information than a bare number like "averaged 432 points."
2. How to Aggregate Returns Across Seeds
A single run (one seed) yields one return curve. When aggregating across seeds, the recommended choice is median + interquartile range (IQR), not mean +/- standard deviation:
| Aggregation | Pros | Cons |
|---|---|---|
| Mean +/- std | Intuitive; everyone can compute it | Skewed by outliers (RL runs occasionally blow up) |
| Median +/- IQR | Robust to outliers; no normality assumption | Slightly less intuitive — but more honest |
| Quantile bands (5/25/75/95%) | The most complete picture | Busier plots |
RL return distributions are usually not normal: once in a while a seed collapses halfway through training (NaN loss, exploration collapse), inflating the variance of the mean. Showing median + IQR for "typical behavior" and full quantile bands for "worst cases" is what responsible reporting looks like.
3. Success Rate and Business Metrics
Return often means nothing to the business. Real decisions are made on success rate, completion rate, cost, and the like. Report both:
| Metric type | Examples | Used for |
|---|---|---|
| Process metrics | Return, TD error, entropy | Debugging and diagnosis |
| Outcome metrics | Task success rate, mean completion time, resource usage | Business decisions |
Balance return with task success rate
Return hides patterns: an average of 400 may come from "80% of episodes score 500 and 20% score 0." Only by reporting success rate alongside return can you distinguish "consistently good" from "gambling good."
4. Evaluation During Training vs. Final Evaluation
- During-training evaluation (every N steps): lightweight — fixed evaluation seeds, 3~5 episodes per seeded environment, for spotting trends only.
- Final evaluation (after training ends): strict — fresh seeds, 10~20 episodes, all randomness disabled (
evalmode,deterministicsampling).
Mixing during-training evaluation with final evaluation
Evaluating every N steps during training and then reporting the last evaluation as the "final score" is a very common way to cheat (usually unintentionally): late in training every evaluation hits the same seed set, the policy may have overfit to those seeds, and the last evaluation may land on a noise peak by chance. The final score must come from an independent final evaluation.
4. Sample Efficiency and Final Performance: Two Dimensions
Every RL evaluation should report two numbers; neither is optional:
text
Final performance (y-axis: level reached)
↑
│ · PPO (high final performance)
│ ·
│ ·
│ ·SAC (sample-efficient)
│ ·
│·
└──────────────────────────→ Sample efficiency (x-axis: steps to reach a level)| Dimension | Definition | How to measure | Why it gets missed |
|---|---|---|---|
| Final performance | Performance when the training budget runs out | Final evaluation | The choice of budget itself influences who looks better |
| Sample efficiency | Steps needed to reach a performance threshold (or AUC) | Read the "first crossing of the threshold" off the learning curve | It's lost the moment you only report final scores |
In engineering you usually care about sample efficiency — your data budget is finite — yet papers and reports often report final performance only. The report template demands both; if one is missing, go back and produce it.
Sample efficiency can also be expressed as AUC (area under the learning curve): integrate the curve, and one number captures both "how fast" and "how high," without depending on a threshold choice. Agarwal et al. (2021) promote it to a core metric alongside median performance.
5. Fairness Pitfalls in Comparisons
Every item below shows up for real, both in paper reproductions and in internal team reviews:
| # | Pitfall | Consequence | Avoidance |
|---|---|---|---|
| 1 | Default hyperparameters for the rival algorithm, tuned ones for yours | A fake "ours is better" | Tune all algorithms equally, or explicitly declare an "official defaults" comparison |
| 2 | Budget counted in epochs while the two algorithms consume steps at different rates | The result is driven by the budget gap | Unify the budget in environment steps |
| 3 | One algorithm tuned to its optimum, the other run casually | Same as above | Fair rounds: tune each independently, then compare at equal budget |
| 4 | Unequal compute (GPU vs CPU, different degrees of parallelism) | Hidden step-count differences | Report wall-clock time and hardware |
| 5 | Different seed sets (A uses 0-4, B uses 100-104) | Comparability is lost | Share one seed set |
| 6 | Randomness left on during evaluation | Extra variance folded into the conclusion | Enable deterministic in final evaluation |
| 7 | Drawing conclusions from a single learning curve | The conclusion rests on one point | Multiple seeds + IQR |
| 8 | Hyperparameters obtained by tuning on the test set | Overfitting to the evaluation set | Reserve holdout environments/tasks (see below) |
Isn't Hyperparameter Tuning Itself a Form of Cheating?
Yes. If you tune repeatedly against the test environment, the score you report is tuned on data that has been seen. Mitigations:
- Reserve a validation environment: tune on set A (environments/seeds), and write the final report on set B (unseen seeds or harder environments).
- Tune sparingly: cap the number of tuning rounds (say, at most 20 per algorithm) and state the tuning budget in the report — reviewers now specifically check this "tuning-budget fairness."
- Report how the defaults do: publish both the "official defaults" and the "tuned" results.
6. Run Management: Experiment Records, Logging, and Tabulation
However rigorous the protocol, sloppy run management will undo it. You need a trio: config records, curve logs, and a results table.
1. Experiment Records
Every run (one (algorithm, seed, environment) combination) should automatically produce a record:
text
experiments/
2026-08-01_ppo_cartpole_s0/
config.yaml # full config (generated by Hydra)
seed.txt # the seed used
train_returns.csv # training return curve
eval_returns.csv # evaluation return curve
metrics.json # final-evaluation summary
model.pt # model weights
git_commit.txt # code version
env_versions.txt # dependency versionsThis is the concrete shape of the "reproducibility hardening" described in Build an RL Project from Scratch. With it, "reproduce a run" simply means "rerun a directory."
2. Logging Tools
| Tool | Strengths | Getting started |
|---|---|---|
| TensorBoard | Scalars, histograms, curves; free and offline | The default choice at the start |
| Weights & Biases (W&B) | Run comparison, tables, collaboration, automatic logging | Essential for team work |
| Hand-rolled CSV | Minimal, zero dependencies | Small single-machine experiments |
| rliable | IQR/quantile-band plotting (statistical aggregation) | Use it for report-stage figures |
Logs should hit disk automatically, not go to the screen
Terminal output gets lost, gets interleaved, and can't be searched. Every metric in the training loop (return, entropy, KL, TD error, learning rate, seed, timestep) should be appended to a CSV or logged automatically via W&B's log(). An unlogged run is a run that never happened.
3. The Master Results Table
When all runs finish, consolidate them into a master table — the core of the evaluation report:
| Algorithm | Environment | Final return (median +/- IQR) | Success rate | Sample efficiency (AUC) | Training steps | Seeds | Hyperparameters |
|---|---|---|---|---|---|---|---|
| PPO | CartPole-v1 | 498 [496, 500] | 100% | 12.4k | 50k | 5 | Tuned |
| SAC | CartPole-v1 | 497 [494, 500] | 100% | 9.8k | 50k | 5 | Tuned |
| ... | ... | ... | ... | ... | ... | ... | ... |
7. Why Offline Evaluation Is So Hard
When all you have is historical data and online trial-and-error is off the table (recommender systems, trading, robot logs), evaluation enters a different league: you can't score a policy by "running the environment," so you're stuck with offline evaluators. The water here runs very deep:
- The behavior policy and the target policy live on different distributions -> trajectory-weight bias (importance-sampling variance explodes with sequence length).
- Offline value estimates run optimistic -> a policy that "looks great" collapses the moment it goes live.
- Bootstrapping errors compound -> value overestimation snowballs (the "bootstrapping hallucination" discussed in Offline RL).
Practical guidance:
| Situation | Trust | Don't trust |
|---|---|---|
| A real online channel exists | A small live A/B test | Pure offline scores (screening only) |
| Only offline data | Conservative offline estimators, cross-checked across methods | Optimistic numbers from a single method |
| Reading a paper | Experiments with real validation | Work that only shows fitted curves |
The "backtest lie" in offline evaluation
Finance calls it backtest overfitting; recommender systems call it offline-metric inflation; machine translation calls it test-set overfitting — same beast: substituting goodness of fit on historical data for future performance. See RL in Financial Trading for an honest dissection. Any "great offline score" conclusion must come with an estimate of how large the unknown uncertainty is.
8. An Evaluation Report Template
Save the template below as EVAL_REPORT.md and fill it in at every project wrap-up. It forces you to pre-empt every point a skeptic would raise:
markdown
# Evaluation Report: <project name> <date>
## 1. Goal and question
- The question this evaluation must answer: ________
- Decision type: ☐ tuning ☐ algorithm comparison ☐ launch
## 2. Experiment design
- Algorithm: ________ (version / code commit: ________)
- Environment: ________ (Gymnasium version: ________)
- Seed set: ________ (is each seed threaded into the environment? ☐ yes ☐ no)
- Training budget: ____ environment steps (unified accounting)
- Hyperparameter source: ☐ official defaults ☐ tuned (tuning rounds: ____; reserved validation environment: ☐ yes ☐ no)
## 3. Metrics and results
| Algorithm | Environment | Final return (median [IQR]) | Success rate | AUC | Seeds |
|---|---|---|---|---|---|
| | | | | | |
- Learning-curve figure (IQR bands): [link]
## 4. Fairness self-check
- [ ] Equal tuning budget for all algorithms / explicit declaration
- [ ] Budget unified in environment steps
- [ ] Final evaluation independent of training evaluation, randomness disabled
- [ ] Software versions and dependencies recorded
## 5. Conclusion and confidence
- Conclusion: ________
- Confidence: ☐ high (multiple seeds + independent evaluation) ☐ medium ☐ low (single run, internal reference only)
- Known limitations: ________Use the same template for internal iterations
You don't need all sections every time — day-to-day tuning only needs an abbreviated version of Sections 2 and 3 (config + curve figure). But every formal, external-facing report (reviews, papers, launch sign-offs) must use the full template. Build the habit from day one; the cost is far lower than reconstructing everything afterward.
9. The Evaluation Workflow End to End
text
Define question → Design matrix → Lock variables → Run experiments → Aggregate curves → Fairness check → Write report → Decide
see Sec 1 see Sec 2 see Secs 2/5 see Sec 6 see Secs 3/4 see Sec 5 see Sec 8In day-to-day tuning, evaluation and tuning are the same loop — every hyperparameter tweak is effectively a mini-evaluation (see Hyperparameter Tuning in Practice). The two share the same logging and curve infrastructure; only the statistical rigor differs.
Further Reading
- Evaluation and Benchmarks — the theory behind this page: why RL evaluation is hard, the benchmark landscape, reporting standards
- Hyperparameter Tuning in Practice — how evaluation protocols and the tuning loop merge; seed pitfalls
- Offline RL — the full expansion of Section 7: why offline evaluation is hard and the conservative fixes
- Build an RL Project from Scratch — wiring the evaluation protocol into a full project pipeline (logging, reproducibility, deployment monitoring)
- Common Pitfalls and Anti-Patterns — a detection checklist for evaluation cheating, seed overfitting, and other failure points
References
- Agarwal, R. et al. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS 2021. arXiv:2108.13264 (the standard reference for robust evaluation protocols — IQR, medians, probability of improvement — with the rliable library)
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
- Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133
- Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292
- rliable library (Google Research): github.com/google-research/rliable (IQR / quantile bands / probability-of-improvement plots)
- Weights & Biases official docs: docs.wandb.ai; TensorBoard: tensorflow.org/tensorboard
- Farama Foundation benchmark notes (maintainers of the Gymnasium/Atari environments): farama.org