Skip to content

Building an RL Evaluation Suite from Scratch

On this page Evaluation protocols in practice — fixed seed matrices, multiple runs, learning curves with IQR, the twin dimensions of sample efficiency vs. final performance; fairness pitfalls in algorithm comparisons; offline and online evaluation.

Building an RL Evaluation Suite from Scratch ​

One-liner: this page teaches you to build an RL evaluation protocol that won't lie to you — from experiment-matrix design and metric selection to fairness pitfalls and the final report, all in one pass. It resolves the recurring puzzle of "training curves keep climbing, yet everything falls apart the moment real comparisons begin," and it serves anyone who has to make decisions from experimental results (tuning hyperparameters, choosing algorithms, deciding whether to ship). By the end, you'll have a reusable evaluation script and a report template of your own.

First, a hard truth: a large share of published RL results cannot be reproduced with a single run. Henderson et al., in "Deep Reinforcement Learning that Matters," put numbers to it: for the same algorithm from the same paper, performance variation across seeds can exceed the variation between different algorithms. Put differently, if you can't evaluate, your experimental conclusions are dice rolls. This page turns dice rolls into readings.

This page is the hands-on companion to Evaluation and Benchmarks: that concept page explains why RL evaluation is hard; this one shows you exactly how to build the evaluation itself.

1. Clarify First: What Question Must This Evaluation Answer ​

An evaluation protocol isn't better for having more pieces — it's better for serving a decision. Start by distinguishing three evaluation goals, because their protocol designs differ completely:

Evaluation goalThe question it must answerKey requirementTypical scope
Hyperparameter decisionDoes changing this hyperparameter help or hurt?Relative comparison is enough; few seeds OKMostly single runs
Algorithm comparisonIs algorithm A really better than B?Multiple seeds + statistical intervals; fairness is unforgivingA multi-seed matrix
Launch decisionCan we ship? What's the cost of failure?Tail-risk analysis, safety boundaries, offline + online validationThe heaviest; long cycles

The most common evaluation mistake: applying a tuning-grade protocol to a launch decision

A "+5 points" bump during tuning may be nothing more than seed luck. Apply the same lax standard to a launch decision, and you've mistaken one draw of variance for a real signal. The more consequential the conclusion, the stricter the evaluation protocol. That's the principle from which every detail on this page follows.

2. Designing the Experiment Matrix: Algorithm x Seed x Environment ​

The basic unit of evaluation is a single run. A run is a quadruple (algorithm, seed, environment, config), and the full set of runs forms the matrix:

text
Experiment matrix (columns = seeds, rows = algorithm x environment)

            seed 0   seed 1   seed 2   seed 3   seed 4
PPO  CartPole    ██████   ██████   ██████   ██████   ██████
PPO  LunarLander ██████   ██████   ██████   ██████   ██████
SAC  CartPole    ██████   ██████   ██████   ██████   ██████
SAC  LunarLander ██████   ██████   ██████   ██████   ██████

█ = one full training run + evaluation (learning curve included)

1. How to Choose Seeds ​

ElementRecommended practiceWhy
Number of seeds>= 5 for comparisons, >= 10 for formal reportsIt takes about 5 seeds for the mean to settle to a distinguishable signal
Seed setFix one set: 0, 1, 2, 3, 4Sharing the same seed set across algorithms maximizes comparability
Seed valuesRecord the full config alongside each seedDifferent seeds and different hyperparameters = uncontrolled variables, no attribution
Environment randomnessThe seed must be threaded into the environment's reset and action_spaceOtherwise the environment is only partially controlled — see Common Pitfalls

2. The Variable-Control Checklist ​

When comparing two algorithms, "algorithm" is the only variable allowed to differ. Everything else must be locked down:

  • [ ] Environment version and wrappers (is TimeLimit on? frame_stack? identical across runs)
  • [ ] Per-seed environment randomness (each seed must seed the environment, not just one global seed at startup)
  • [ ] Observation preprocessing (normalization, grayscale, frame stacking)
  • [ ] Total training-step budget (not episodes or iterations — what an "epoch" means varies by algorithm, so the budget must be counted in environment interactions)
  • [ ] The evaluation protocol itself (see Section 3)
  • [ ] Software versions (pin the exact gymnasium/numpy/torch versions in the report)

Why the training budget must be counted in environment steps, not training iterations

One PPO update consumes 2,048 steps; one REINFORCE update consumes a whole episode (tens to hundreds of steps). Align budgets by "number of updates" and PPO gets a massive, invisible advantage. Align them by environment steps, and both sides spend the same money to buy the same amount of data — a fair comparison. This is also the premise for sample efficiency vs. final performance.

3. Metrics: Mean Return, Learning Curves, Success Rate ​

1. The Learning Curve Is the Mother of All Metrics ​

Don't just report a "final score." The learning curve (x-axis: environment steps, y-axis: evaluation return) preserves the time dimension: who converges fast, who's degrading, who's oscillating early on. A single learning-curve plot with an IQR band across seeds carries an order of magnitude more information than a bare number like "averaged 432 points."

2. How to Aggregate Returns Across Seeds ​

A single run (one seed) yields one return curve. When aggregating across seeds, the recommended choice is median + interquartile range (IQR), not mean +/- standard deviation:

AggregationProsCons
Mean +/- stdIntuitive; everyone can compute itSkewed by outliers (RL runs occasionally blow up)
Median +/- IQRRobust to outliers; no normality assumptionSlightly less intuitive — but more honest
Quantile bands (5/25/75/95%)The most complete pictureBusier plots

RL return distributions are usually not normal: once in a while a seed collapses halfway through training (NaN loss, exploration collapse), inflating the variance of the mean. Showing median + IQR for "typical behavior" and full quantile bands for "worst cases" is what responsible reporting looks like.

3. Success Rate and Business Metrics ​

Return often means nothing to the business. Real decisions are made on success rate, completion rate, cost, and the like. Report both:

Metric typeExamplesUsed for
Process metricsReturn, TD error, entropyDebugging and diagnosis
Outcome metricsTask success rate, mean completion time, resource usageBusiness decisions

Balance return with task success rate

Return hides patterns: an average of 400 may come from "80% of episodes score 500 and 20% score 0." Only by reporting success rate alongside return can you distinguish "consistently good" from "gambling good."

4. Evaluation During Training vs. Final Evaluation ​

  • During-training evaluation (every N steps): lightweight — fixed evaluation seeds, 3~5 episodes per seeded environment, for spotting trends only.
  • Final evaluation (after training ends): strict — fresh seeds, 10~20 episodes, all randomness disabled (eval mode, deterministic sampling).

Mixing during-training evaluation with final evaluation

Evaluating every N steps during training and then reporting the last evaluation as the "final score" is a very common way to cheat (usually unintentionally): late in training every evaluation hits the same seed set, the policy may have overfit to those seeds, and the last evaluation may land on a noise peak by chance. The final score must come from an independent final evaluation.

4. Sample Efficiency and Final Performance: Two Dimensions ​

Every RL evaluation should report two numbers; neither is optional:

text
        Final performance (y-axis: level reached)
             ↑
             │        · PPO (high final performance)
             │        ·
             │        ·
             │   ·SAC (sample-efficient)
             │   ·
             │·
             └──────────────────────────→ Sample efficiency (x-axis: steps to reach a level)
DimensionDefinitionHow to measureWhy it gets missed
Final performancePerformance when the training budget runs outFinal evaluationThe choice of budget itself influences who looks better
Sample efficiencySteps needed to reach a performance threshold (or AUC)Read the "first crossing of the threshold" off the learning curveIt's lost the moment you only report final scores

In engineering you usually care about sample efficiency — your data budget is finite — yet papers and reports often report final performance only. The report template demands both; if one is missing, go back and produce it.

Sample efficiency can also be expressed as AUC (area under the learning curve): integrate the curve, and one number captures both "how fast" and "how high," without depending on a threshold choice. Agarwal et al. (2021) promote it to a core metric alongside median performance.

5. Fairness Pitfalls in Comparisons ​

Every item below shows up for real, both in paper reproductions and in internal team reviews:

#PitfallConsequenceAvoidance
1Default hyperparameters for the rival algorithm, tuned ones for yoursA fake "ours is better"Tune all algorithms equally, or explicitly declare an "official defaults" comparison
2Budget counted in epochs while the two algorithms consume steps at different ratesThe result is driven by the budget gapUnify the budget in environment steps
3One algorithm tuned to its optimum, the other run casuallySame as aboveFair rounds: tune each independently, then compare at equal budget
4Unequal compute (GPU vs CPU, different degrees of parallelism)Hidden step-count differencesReport wall-clock time and hardware
5Different seed sets (A uses 0-4, B uses 100-104)Comparability is lostShare one seed set
6Randomness left on during evaluationExtra variance folded into the conclusionEnable deterministic in final evaluation
7Drawing conclusions from a single learning curveThe conclusion rests on one pointMultiple seeds + IQR
8Hyperparameters obtained by tuning on the test setOverfitting to the evaluation setReserve holdout environments/tasks (see below)

Isn't Hyperparameter Tuning Itself a Form of Cheating? ​

Yes. If you tune repeatedly against the test environment, the score you report is tuned on data that has been seen. Mitigations:

  1. Reserve a validation environment: tune on set A (environments/seeds), and write the final report on set B (unseen seeds or harder environments).
  2. Tune sparingly: cap the number of tuning rounds (say, at most 20 per algorithm) and state the tuning budget in the report — reviewers now specifically check this "tuning-budget fairness."
  3. Report how the defaults do: publish both the "official defaults" and the "tuned" results.

6. Run Management: Experiment Records, Logging, and Tabulation ​

However rigorous the protocol, sloppy run management will undo it. You need a trio: config records, curve logs, and a results table.

1. Experiment Records ​

Every run (one (algorithm, seed, environment) combination) should automatically produce a record:

text
experiments/
  2026-08-01_ppo_cartpole_s0/
    config.yaml        # full config (generated by Hydra)
    seed.txt           # the seed used
    train_returns.csv  # training return curve
    eval_returns.csv   # evaluation return curve
    metrics.json       # final-evaluation summary
    model.pt           # model weights
    git_commit.txt     # code version
    env_versions.txt   # dependency versions

This is the concrete shape of the "reproducibility hardening" described in Build an RL Project from Scratch. With it, "reproduce a run" simply means "rerun a directory."

2. Logging Tools ​

ToolStrengthsGetting started
TensorBoardScalars, histograms, curves; free and offlineThe default choice at the start
Weights & Biases (W&B)Run comparison, tables, collaboration, automatic loggingEssential for team work
Hand-rolled CSVMinimal, zero dependenciesSmall single-machine experiments
rliableIQR/quantile-band plotting (statistical aggregation)Use it for report-stage figures

Logs should hit disk automatically, not go to the screen

Terminal output gets lost, gets interleaved, and can't be searched. Every metric in the training loop (return, entropy, KL, TD error, learning rate, seed, timestep) should be appended to a CSV or logged automatically via W&B's log(). An unlogged run is a run that never happened.

3. The Master Results Table ​

When all runs finish, consolidate them into a master table — the core of the evaluation report:

AlgorithmEnvironmentFinal return (median +/- IQR)Success rateSample efficiency (AUC)Training stepsSeedsHyperparameters
PPOCartPole-v1498 [496, 500]100%12.4k50k5Tuned
SACCartPole-v1497 [494, 500]100%9.8k50k5Tuned
........................

7. Why Offline Evaluation Is So Hard ​

When all you have is historical data and online trial-and-error is off the table (recommender systems, trading, robot logs), evaluation enters a different league: you can't score a policy by "running the environment," so you're stuck with offline evaluators. The water here runs very deep:

  • The behavior policy and the target policy live on different distributions -> trajectory-weight bias (importance-sampling variance explodes with sequence length).
  • Offline value estimates run optimistic -> a policy that "looks great" collapses the moment it goes live.
  • Bootstrapping errors compound -> value overestimation snowballs (the "bootstrapping hallucination" discussed in Offline RL).

Practical guidance:

SituationTrustDon't trust
A real online channel existsA small live A/B testPure offline scores (screening only)
Only offline dataConservative offline estimators, cross-checked across methodsOptimistic numbers from a single method
Reading a paperExperiments with real validationWork that only shows fitted curves

The "backtest lie" in offline evaluation

Finance calls it backtest overfitting; recommender systems call it offline-metric inflation; machine translation calls it test-set overfitting — same beast: substituting goodness of fit on historical data for future performance. See RL in Financial Trading for an honest dissection. Any "great offline score" conclusion must come with an estimate of how large the unknown uncertainty is.

8. An Evaluation Report Template ​

Save the template below as EVAL_REPORT.md and fill it in at every project wrap-up. It forces you to pre-empt every point a skeptic would raise:

markdown
# Evaluation Report: <project name> <date>

## 1. Goal and question
- The question this evaluation must answer: ________
- Decision type: ☐ tuning ☐ algorithm comparison ☐ launch

## 2. Experiment design
- Algorithm: ________ (version / code commit: ________)
- Environment: ________ (Gymnasium version: ________)
- Seed set: ________ (is each seed threaded into the environment? ☐ yes ☐ no)
- Training budget: ____ environment steps (unified accounting)
- Hyperparameter source: ☐ official defaults ☐ tuned (tuning rounds: ____; reserved validation environment: ☐ yes ☐ no)

## 3. Metrics and results
| Algorithm | Environment | Final return (median [IQR]) | Success rate | AUC | Seeds |
|---|---|---|---|---|---|
|      |      |                     |        |     |         |

- Learning-curve figure (IQR bands): [link]

## 4. Fairness self-check
- [ ] Equal tuning budget for all algorithms / explicit declaration
- [ ] Budget unified in environment steps
- [ ] Final evaluation independent of training evaluation, randomness disabled
- [ ] Software versions and dependencies recorded

## 5. Conclusion and confidence
- Conclusion: ________
- Confidence: ☐ high (multiple seeds + independent evaluation) ☐ medium ☐ low (single run, internal reference only)
- Known limitations: ________

Use the same template for internal iterations

You don't need all sections every time — day-to-day tuning only needs an abbreviated version of Sections 2 and 3 (config + curve figure). But every formal, external-facing report (reviews, papers, launch sign-offs) must use the full template. Build the habit from day one; the cost is far lower than reconstructing everything afterward.

9. The Evaluation Workflow End to End ​

text
Define question → Design matrix → Lock variables → Run experiments → Aggregate curves → Fairness check → Write report → Decide
  see Sec 1        see Sec 2      see Secs 2/5      see Sec 6       see Secs 3/4     see Sec 5     see Sec 8

In day-to-day tuning, evaluation and tuning are the same loop — every hyperparameter tweak is effectively a mini-evaluation (see Hyperparameter Tuning in Practice). The two share the same logging and curve infrastructure; only the statistical rigor differs.

Further Reading ​

References ​

  • Agarwal, R. et al. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS 2021. arXiv:2108.13264 (the standard reference for robust evaluation protocols — IQR, medians, probability of improvement — with the rliable library)
  • Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
  • Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133
  • Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292
  • rliable library (Google Research): github.com/google-research/rliable (IQR / quantile bands / probability-of-improvement plots)
  • Weights & Biases official docs: docs.wandb.ai; TensorBoard: tensorflow.org/tensorboard
  • Farama Foundation benchmark notes (maintainers of the Gymnasium/Atari environments): farama.org