Appearance
Common Pitfalls and Anti-Patterns
In a nutshell: this page collects the ten most common ways RL projects blow up — symptom, root cause, and detection method in one place — plus a diagnostic checklist for "I tried to reproduce a paper and failed." It serves as the troubleshooting entry point for the three classic frustrations: "training won't improve," "the results aren't trustworthy," and "I can't reproduce this." By the end you'll have a roadmap for what to check first, second, and third.
Debugging RL is fundamentally harder than debugging supervised learning for one reason: gradient descent in supervised learning has a visible target, whereas everything in RL passes through layers of amplification — environment, policy, value estimation — so a bug in any single stage can masquerade as "the algorithm doesn't work." This page lays out the most common traps in a "symptom → root cause → detection" structure.
1. The Pitfalls at a Glance
| # | Pitfall | Symptom in one sentence | Damage | How to detect |
|---|---|---|---|---|
| 1 | Reward hacking | Return skyrockets while behavior gets increasingly absurd | ★★★★★ | Behavior audit + per-term reward attribution |
| 2 | Seed overfitting | Change the seed and the magic disappears | ★★★★ | Multi-seed validation on a fresh seed set |
| 3 | Silent environment bugs | Training looks stable but the results are meaningless | ★★★★★ | Environment acceptance checklist + smoke tests |
| 4 | Evaluation cheating | The score comes from the training distribution | ★★★★ | Independent final evaluation |
| 5 | Sample-efficiency illusion | "It improved" — but only by burning 10x the compute | ★★★ | Report steps + wall-clock time |
| 6 | Curse of dimensionality | The state space is too large to learn | ★★★★ | Feature analysis, dimensionality reduction / embeddings |
| 7 | Lost logs | You can't reconstruct where a result came from | ★★★★ | Automated logging, experiment directories |
| 8 | Parallelization bugs | Works fine single-process, breaks multi-process | ★★★★ | Parallel consistency tests |
| 9 | Environment mismatch | Training and deployment environments differ | ★★★ | Environment validation + version pinning |
| 10 | Implementation detail gaps | "Reproducing the paper" always falls slightly short | ★★★★★ | Line-by-line comparison with the official implementation |
Each is expanded below. Reward hacking, evaluation, and seed issues get deeper treatment on other pages of this site; here they're covered in a condensed "engineering perspective."
2. Pitfall 1: Reward Hacking and Specification Gaming
- Symptom: The learning curve looks beautiful, but the policy's behavior is clearly "cheating" — taking shortcuts, farming points, and exploiting loopholes in the spec.
- Root cause: The reward function is an approximation of your true objective. The agent optimizes the reward, not your intent. This is the central concern of reward engineering.
- Greatest hits:
- A boat sailing in circles inside a lagoon to farm points, because "visiting new places" earns reward;
- A robot learning to block its camera to fake "task complete";
- A game AI wedging itself into a bugged position to score infinitely;
- In RLHF, a model learning to say "I can't answer that" to dodge punishment for low-quality responses (a form of alignment failure — see LLM alignment).
- Detection:
- Behavior audit: Visualize or sample-replay the policy's rollouts, and sanity-check that the behavior looks reasonable to a human eye.
- Reward attribution: Break reward down by term and confirm the policy is earning the parts you want, not the parts you don't.
- Adversarial testing: Deliberately construct inputs that expose reward loopholes and see whether the policy takes the bait.
The most dangerous signal: return up, business metrics flat
If return climbs 50% while business metrics (user retention, task success rate) don't budge, you're probably looking at reward hacking or a distorted reward scale — don't celebrate, run a behavior audit first. The financial-domain version of this trap (fitting the backtest instead of achieving real performance) is covered in RL in finance.
3. Pitfall 2: Seed Overfitting
- Symptom: A configuration looks amazing at
seed=0and falls apart atseed=1; or your "hyperparameter tuning" has merely memorized one set of random numbers. - Root cause: RL results have huge variance (especially for on-policy and policy-gradient methods), so good news from a single seed is very likely just noise. Tuning repeatedly against test seeds = overfitting to noise.
- Detection:
- Run 3–5 seeds per configuration and look at the median and IQR (the spread), not single points.
- Keep the seed set used for tuning completely separate from the seed set used for final validation.
text
The single-seed illusion (same configuration, three seeds):
seed 0: ──────────── 480 points
seed 1: ────╱╲────── 210 points
seed 2: ──────────╱╲ 310 points
Only seed 0 → "Tuning success!"
All seeds → "Just luck."- Fix: Nail down an evaluation protocol (multi-seed + IQR); see evaluation in practice and tuning in practice.
4. Pitfall 3: Environment Bugs Silently Corrupting Learning
- Symptom: Training runs fine with no errors, but the results are untrustworthy — the reward might be computed wrong,
donetriggered wrong, observations written wrong, or randomness left unseeded. - Root cause: Environment bugs don't throw exceptions. They silently contaminate the learning signal while everything looks like it's "training."
- Classic cases:
- Confusing
terminatedwithtruncated: treating "survived 500 steps" in CartPole as failure completely inverts the learning signal (see the pitfall table in the tutorial page). - Using the wrong variable in the reward formula (computing "closeness" from position but with the sign flipped).
resetnot actually resetting internal state (stale state leaks through).- Observations missing dimensions or in the wrong order, so the policy "sees" something different from what you think it sees.
- Confusing
- Detection: The environment acceptance checklist is the only reliable defense (full version in build your own RL project from scratch):
- A 1,000-step smoke test with a random policy;
- A hand-written heuristic to confirm the environment is "winnable";
- Two rollouts with the same seed producing identical results;
- Printing every observation dimension and reward individually for manual inspection.
Three clues that say "suspect the environment first"
Training won't improve; the return magnitude is abnormal (too small / too large / suddenly NaN); or the same code behaves completely differently in a different environment. All three point at the environment — run the acceptance checklist before touching the algorithm.
5. Pitfall 4: Evaluation Cheating (Testing on the Training Distribution)
- Symptom: Reported scores are far above true performance; everything shrinks dramatically after deployment.
- Root cause: Evaluation shares the same seeds / environments / randomness as training, so the evaluation "remembers" states seen during training; or randomness is left on during evaluation; or the final training curve is simply presented as the final score.
- Detection:
- Final evaluation uses fresh seeds, deterministic action selection (
deterministic), and an evaluation set fully separated from the training seeds. - Record the environment seed for every evaluation and spot-check for overlap with training.
- Present learning curves and final scores separately.
- Final evaluation uses fresh seeds, deterministic action selection (
Full details of the evaluation protocol are in build an RL evaluation setup from scratch; the theoretical version of "why evaluation is so hard" is in evaluation and benchmarks.
Unintentional cheating is the most common kind
Most evaluation cheating isn't deliberate: the evaluation environment forgot the seed, or the EvalCallback in training reused the training seeds. Write the evaluation protocol into code instead of keeping it in your head — that's the only reliable defense.
6. Pitfall 5: The Sample-Efficiency Illusion (Gains Bought with Compute)
- Symptom: The model "improved" — but look closer and it's the product of 10x the samples/compute. The same effect was reached more cheaply by the baseline.
- Root cause: Only final performance is considered, not sample efficiency; unequal budgets let different algorithms "eat" unequal amounts of the problem.
- Detection:
- Reports must include environment interaction steps (not epochs or gradient updates).
- Compare at a unified step budget; report wall-clock time and hardware alongside.
- Use learning curves to show "steps required to reach a given performance" (AUC), not just the final score.
Before switching algorithms, ask: does the sample budget support it?
On-policy methods (PPO) use each sample once; off-policy methods (SAC) reuse them. If your application can only collect 50k steps, PPO may not even complete one decent learning cycle — this is exactly the "sample-efficiency awareness" rule from design principles blowing up in practice.
7. Pitfall 6: The Curse of Dimensionality and Missing Feature Engineering
- Symptom: Once state dimensionality gets high (say, 100 raw sensor dimensions, or raw pixels fed into a tabular method), the algorithm can't learn, or learns absurdly slowly.
- Root cause: The number of grid cells in tabular methods explodes exponentially with dimension; deep RL also needs an appropriate representation and normalization for high-dimensional states.
- Detection:
- Do observation dimensions differ in scale by huge factors (1,000x or more)? → Normalization needed.
- Does the state mix in noise dimensions irrelevant to the decision? → Do feature selection / embeddings.
- Are you using tabular methods or linear approximation on a nonlinear, high-dimensional problem? → Switch to a deep representation.
| Symptom | Treatment |
|---|---|
| Scale differences of 1,000x | Observation normalization (running normalization) |
| High-dimensional + continuous | Deep network + feature embeddings |
| Tabular explosion | Switch to function approximation, or restructure the state representation |
| Insufficient information | Frame stacking / memory (RNN/Transformer) |
8. Pitfalls 7, 8, and 9: The Engineering Trio
These three are grouped together because they're "basic engineering hygiene" problems — they look trivial but devour a huge amount of RL teams' time.
Pitfall 7: Lost Logs
- Symptom: You want to look up the hyperparameters behind "that good run from last time" — and discover nothing was recorded.
- Fix: Automatically persist every training-loop metric (CSV/W&B), with configuration and results living in the same experiment directory. The principle lives in section 6 of evaluation in practice.
Pitfall 8: Parallelization Bugs
- Symptom: Runs great with a single environment; the moment
SubprocVecEnv/multi-processing kicks in, results degrade or crash. - Root cause: The parallel environments weren't seeded independently (multiple subprocesses sharing the same random state); or the observation array is an alias into shared memory and gets modified in parallel.
- Detection:
- With N parallel environments, print all N observations after
resetand confirm they're all different. - Run once single-process and once multi-process; compare whether "average return at the same seed" is in the same ballpark.
- Use
copy.deepcopyfor isolation, or avoid in-place modification of observations.
- With N parallel environments, print all N observations after
Pitfall 9: Training/Deployment Environment Mismatch
- Symptom: A trained model plunges in performance when moved to a different environment (Python version, GPU, dependencies).
- Root cause: RL is highly sensitive to versions (gymnasium/numpy/torch versions can all change behavior).
- Fix: Pin dependencies (lockfile/Docker), record
pip freeze+ git commit + machine info, and run a "smoke regression" with the same image before deployment.
9. Diagnostic Checklist for "I Failed to Reproduce a Paper"
Failing to reproduce an RL paper happens to everyone. Before blaming yourself or the paper, work through this checklist in order of decreasing probability:
text
□ 1. Official-implementation detail gaps
Paper formulas vs. official code: is reward normalization done? Which GAE variant?
How is advantage normalized? → Reproduce with the official code first, then port to yours
□ 2. Hyperparameter mismatch
Appendix hyperparameters vs. main-text tables? Environment version off? → Compare field by field
□ 3. Environment version differences
CartPole-v0 vs. v1? Atari no-op/frame-stacking parameters? → Match the paper's environment config
□ 4. Seeds and randomness
Did the paper seed runs? Best/mean/median over how many seeds? → Align the statistical protocol
□ 5. Training budget
Paper used 100M steps, you ran 1M? → Check you're in the same ballpark
□ 6. Evaluation protocol
Does the paper's "score" come from the last training evaluation or an independent one? → Align it
□ 7. A genuine bug in your implementation
If everything above checks out → Compare line by line against the official implementation,
focusing on done handling, gradient clipping, LR scheduling, and exploration annealingWhy the "official implementation" is the first place to look
Engstrom et al., in Implementation Matters in Deep RL, showed that PPO's advantage over TRPO comes almost entirely from implementation details rather than the algorithmic differences in the paper. So when reproduction fails, check the implementation before doubting the theory — a lesson the RL replication community has learned the hard way.
10. A Troubleshooting Roadmap: Stringing the Ten Pitfalls Together
When results look wrong, work through this sequence (change one thing at a time, keeping everything else fixed):
text
Training not improving / results look suspicious?
│
├─① Is the environment trustworthy? → Environment acceptance checklist (Pitfall 3)
├─② Is the baseline right? → Run random + heuristic policies (to see where the bar sits)
├─③ Is this a single-seed illusion? → Multi-seed validation (Pitfall 2)
├─④ Is the reward deceiving me? → Behavior audit + reward attribution (Pitfall 1)
├─⑤ Is evaluation cheating? → Independent final evaluation (Pitfall 4)
├─⑥ Is compute propping this up? → Look at steps vs. returns (Pitfall 5)
├─⑦ Is the state representation right? → Normalization / embeddings (Pitfall 6)
├─⑧ Is my engineering hygiene okay? → Logs / parallelism / versions (Pitfalls 7–9)
└─⑨ Are the implementation details right? → Compare against the official implementation (Pitfall 10)11. The Everyday Debugging Workbench: Making Troubleshooting Routine
The first ten sections cover "troubleshooting after something goes wrong," but the better deal is to weld diagnostics into every training run, so most pits are visible before they grow into disasters. Below is a minimal workbench that "records while training, so you can inspect when things break":
1. Diagnostics You Must Log on Every Run
| Metric | Log frequency | What pitfall it exposes |
|---|---|---|
| Train return / eval return | Every N steps | The first crime scene for every problem |
| Policy entropy | Every update | Exploration collapse (entropy plunges), policy randomization (entropy won't drop) |
| Value loss / TD error | Every update | Learning rate too high, unstable bootstrapping |
| Gradient norm | Every update | Gradient explosion (the prelude to NaN) |
| Parameter norm after updates | Every update | Parameter drift (learning rate too high) |
| Q value / advantage mean | Every update | Value overestimation drift |
| Per-term reward breakdown | Every episode | Which reward term dominates (the attribution behind Pitfall 1) |
| Environment seed + version | Every experiment | Reproducibility (Pitfall 9) |
2. Three "Alarms"
Write thresholds into the training script — alarms are far faster than post-mortem log digging:
python
# Three "watchdogs" inside the training loop (illustrative)
if entropy < 0.1: # Entropy too low: exploration collapse
warn("entropy collapsed, check exploration / ent_coef")
if torch.isnan(policy_loss): # NaN: gradient explosion or NaN reward
warn("NaN detected, dump config and halt")
if eval_return < random_baseline: # Below random baseline: suspect environment/reward
warn("below random baseline after 10k steps, check reward signal")3. A "Pitfall Knowledge Base" Template
Distill the team's failures into searchable entries (this is the concrete format for the "document your failures" rule in design principles):
markdown
# KNOWN_FAILURES.md — team pitfall knowledge base
## [2026-08] PPO intermittently goes NaN on GPU, fine on CPU
- Symptom: loss turns NaN after 2k steps; cannot reproduce on CPU
- Root cause: running stats for reward normalization not synchronized across
processes (each process computed its own)
- Detection: compare reward mean/std for the first 500 steps on CPU vs. GPU
- Fix: maintain normalization statistics in the main process only
- Lesson: when GPU and CPU disagree, first ask "who owns the global state"Add an entry every time you burn more than a day on a bug — past about 20 entries, the team is essentially immune to these pits.
4. Debugging Tools Cheat Sheet
| Scenario | Tool | How to use it |
|---|---|---|
| Watching training scalar curves | TensorBoard / W&B | Log every diagnostic metric and compare runs |
| Seeing what rollouts look like | render_mode="human" or save a GIF | Behavior audit (Pitfall 1) |
| Reproducing a single seed | Command-line --seed N | Hierarchical seeding + config archiving |
| Comparing multi-seed distributions | rliable | IQR bands, probability-of-improvement plots |
| Breakpoint-debugging the environment | A standalone test_env.py | Print obs/reward/done at every step and eyeball them |
The gold standard for a debugging workbench
A newcomer taking over the project should be able to "look at the diagnostic metrics in the logs and say what state the training is in." If the diagnostics are incomplete or nobody reads them, your project still troubleshoots by "run it and look at the score" — which is driving with your eyes closed.
Further Reading
- Reward engineering — the full expansion of Pitfall 1: a case collection of reward hacking and anti-gaming design
- Evaluation and benchmarks — the theory behind Pitfalls 4 and 5: why RL evaluation is hard and the benchmark landscape
- RL in finance — the extreme form of Pitfalls 1/5 in finance: the lies of backtesting and sample scarcity
- Build an RL evaluation setup from scratch — the complete protocol for Pitfalls 2 and 4: multi-seed, IQR, fairness
- Tuning and hyperparameter optimization — the engineering side of Pitfalls 2 and 10: diagnose before you tune, check the implementation before reproducing
References
- Amodei, D. et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565 (the canonical discussion of reward misspecification and exploration risk)
- Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
- Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133
- Agarwal, R. et al. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS 2021. arXiv:2108.13264
- Gymnasium documentation on terminated/truncated (Farama Foundation): gymnasium.farama.org
- Irpan, A. (2018). Deep Reinforcement Learning Doesn't Work Yet. alexirpan.com/2018/02/14/rl-hard.html