Appearance
RL Design Principles
One-liner: this page distills the lessons of repeatedly failed RL projects into ten design principles — each paired with a real-world counterexample and one actionable practice. It addresses the root cause of "why is my RL project always stuck in overtime rework," and it serves tech leads, solo developers, and anyone who has to actually ship an RL system. By the end, you'll walk away with two pieces of paper: the ten principles to pin above your desk, and a project self-check checklist.
Start with the facts. Why do RL projects fail significantly more often than ordinary ML projects? Because RL failure modes are more numerous, better hidden, and surface later:
- Failures hide better: when training stalls, you can't tell whether it's the algorithm, the hyperparameters, the reward, or an environment bug (Common Pitfalls covers this in full).
- Failures arrive late: supervised learning reveals bad data within hours; RL often runs for a day or two before you discover it "learned the wrong signal."
- Failures cost more: trial and error needs an environment, and the more realistic the environment, the pricier — robots, trading, and real users all carry enormous trial-and-error costs.
These ten principles grew out of exactly those failures. They aren't inspiration — they're a memo against paying the same tuition twice.
1. Why RL Projects Fail: Three "Failure Bills"
| Failure type | Typical share (informal estimate) | Representative symptoms | Related principles |
|---|---|---|---|
| Wrong problem formulation | ~30% | "This isn't an RL problem at all," or a mis-specified reward | Principles 1, 2 |
| Engineering problems | ~30% | Environment bugs, irreproducible results, lost logs | Principles 2, 5, 7 |
| Algorithm / tuning problems | ~25% | Won't learn, collapses, high variance | Principles 3, 4, 6 |
| Business / deployment problems | ~15% | Built but unused; metrics don't match the business | Principles 8, 9, 10 |
Note: only a quarter of the problems are "algorithm" problems. Yet the vast majority of team resources get poured into algorithms — which is precisely why Principle 1 exists.
2. The Ten Principles at a Glance
| # | Principle | In one line | The counterexample |
|---|---|---|---|
| 1 | Reward before algorithm | The reward is the product spec — write the reward doc before picking an algorithm | A mis-specified reward wastes even the best PPO |
| 2 | Environment first | The environment is the product; its quality decides everything | One environment bug wastes two weeks for everyone |
| 3 | Simplicity first | Start with the simplest viable strategy and upgrade step by step | Reach for SAC on day one; three days of silence |
| 4 | Baselines first | Without a baseline, "getting better" is unverifiable | No random/heuristic baseline -> a false "the algorithm works" |
| 5 | Reproducible experiments | An experiment you can't reproduce is one you didn't run | Results flip completely on a different machine |
| 6 | Safety guardrails | Decide what failure looks like before going live | The policy flails and production goes down |
| 7 | Sample-efficiency awareness | Every interaction has a cost — budget first | A 10k-step training budget with an on-policy algorithm |
| 8 | Small-step iteration | Change one thing at a time, one cycle a day | Five variables changed at once; two hours later, no attribution |
| 9 | Document failures | Turn every failure into a lessons library | The same pitfall gets stepped into by three separate teams |
| 10 | Align with the business | Metrics must translate into business language | Great algorithm scores, unmoved business metrics |
3. Principle 1: Design the Reward Before Choosing the Algorithm
Counterexample: A team designed a reward to "stop NPCs from getting stuck on walls" — "+1 for every step closer to the goal, -10 for hitting a wall." The agent learned to hug the wall (+1 every step is a sure thing; the wall-collision probability is low). The better the algorithm optimized, the more "cleverly bad" the behavior became.
The reward is the only machine-readable expression of the product requirements. Your choice of algorithm only determines whether learning happens at all; the reward determines what the learned thing looks like.
Actionable Practices
- Before writing any code, draft a one-page reward-design document: the source, value, and rationale of each reward term, plus the loopholes it might invite.
- Health-check the reward magnitude: run a random policy and look at the average per-step reward — if it's on the order of 1e-4, floating-point precision alone is polluting learning; normalize first.
- Run a "reward audit" before launch: enumerate the cheapest score-farming paths the policy might find, and close them one by one.
Links to Related Pages
The full methodology of reward design (potential-based shaping, sparse-reward countermeasures, a case collection of reward hacking, inverse RL) lives in Reward Engineering. In RLHF settings the "reward" becomes a model of human preferences, and this principle applies just as much — over-optimizing the reward model is precisely an alignment failure; see LLM Alignment: RLHF in Practice.
The one-sentence test
Explain your reward function to a non-technical person. If they can point out "how the agent could cheat for points," your reward still has holes.
4. Principle 2: Environment First (The Environment Is the Product)
Counterexample: A scheduling project trained in a home-grown simulator for four weeks; the agent performed flawlessly. At launch they discovered the simulator's "machine failure probability" was hard-coded at 5%, while the real production line ran at 15% and drifted over time — the policy collapsed on the true distribution.
The line "the environment is the product" in Anatomy of an RL System is not a slogan: the environment determines your data distribution, and the data distribution determines what the policy can learn. Get the environment wrong, and everything after is wasted effort.
Actionable Practices
- Environment acceptance checklist (the full list is in Build an RL Project from Scratch): smoke tests, a random policy running 1,000 steps, bounded rewards, same-seed reproducibility, a reachable heuristic.
- Verify "the environment is trustworthy" before writing any algorithm: run the simplest possible policy to confirm a learning signal exists before spending algorithm effort.
- Decouple the environment from training: keep it as a standalone service/module so swapping implementations (Python -> C++ -> distributed) never touches training code.
- Domain-randomization awareness: the simulator and reality have a gap, so randomize physics parameters during training (robotics scenarios: see Robot Control and Sim2Real).
The environment is also the requirements spec
Write the environment like a product requirements document: what the states are, action bounds, the reward formula, termination conditions, sources of randomness. That document doubles as test cases, a review artifact, and handoff material — it deserves more review rounds than the algorithm code.
5. Principle 3: Start Simple (Random -> Heuristic -> Simple RL -> Complex RL)
Counterexample: A team reached for SAC the moment the project was approved, with 8 GPUs running parallel tuning. Two weeks later, the return hadn't budged. They switched to the plainest ε-greedy linear policy, which caught up with the heuristic baseline in a single day — the problem had never needed SAC's complexity in the first place.
Complexity is a liability: the more complex the algorithm, the harder to debug, reproduce, and explain. Run the simplest thing first, until simplicity is proven insufficient.
The Complexity Ladder
text
Random policy ──→ Heuristic/rules ──→ Simple RL (DQN/REINFORCE)
│ │ │
Verify the env Verify the problem Verify the RL signal path
is interactive is solvable │
Complex RL (PPO/SAC) + tuning
│
Upgrade only when simple falls short| Rung | Purpose | When to move down |
|---|---|---|
| Random | Environment smoke test, reward collection | Immediately |
| Heuristic | Problem solvability, reward sanity | Once it scores |
| Simple RL | RL signal path, correct training pipeline | Once the pipeline is verified |
| Complex RL | Push the ceiling | Once simple is proven insufficient |
6. Principle 4: Baselines First
Counterexample: The team tuned PPO for over a month; the final return was 3x the random policy, and everyone celebrated. Then they ran the heuristic baseline: a 50-line greedy rule achieved 4.5x random. Three months of "algorithm work" had added nothing.
"Baselines first" is the cheapest, most life-saving principle: it turns "is my algorithm any good?" into an answerable question.
Actionable Practices
- Within the first week, produce complete learning curves for a random baseline and a heuristic baseline.
- Overlay every subsequent experiment on these two lines (two
plt.axhlinereference lines are enough). - Keep the evaluation protocol identical to the baselines' (same seeds, budget, metrics) — otherwise the baseline is meaningless. Protocol details: Building an RL Evaluation Suite from Scratch.
Don't make baselines a formality
If "baselines first" happens once and is never updated, it never happened. The baseline's value lies in being a standing reference: every algorithm change must answer "where is it better than the baseline, and by how much?" Keep the baseline lines on every plot until the project ends.
7. Principle 5: Reproducible Experiments
Counterexample: Engineer A ran PPO with PyTorch 1.13 and got 480 points; Engineer B reproduced the same code in a fresh environment with PyTorch 2.0 and got 390. After three days of arguing, they found the cause: the
numpyrandom state had never been seeded inside the environment — the same seed followed different environment sequences on the two machines.
Reproducibility in RL is worse than in supervised learning: same seed, same code — change the GPU, the library version, or the parallelism scheme, and the results change.
Actionable Practices
- Three-layer seeding: global (python/numpy/torch) + environment (
reset(seed=...)) + algorithm (framework seed parameter). - Config as code: experiment configs live in YAML/Hydra, can be overridden from the command line, and are archived automatically with every run.
- Record everything: git commit, dependency versions, and machine info go into the experiment directory.
- Multiple seeds are part of reproducibility: what you reproduce is the "performance distribution," not "one curve."
Full engineering details are in the reproducibility-hardening section of Build an RL Project from Scratch; the diagnostic checklist for "my paper reproduction failed" is in Common Pitfalls.
8. Principle 6: Safety Guardrails
Counterexample: A ranking policy for recommendations looked good in offline evaluation only, with no shadow mode. On its first live day, the policy took extreme actions for a small slice of users (recommending the same product over and over), CTR cratered, and the rollback took three hours.
RL policies are more likely than rule-based systems to "cleverly cross the line." Guardrails are not optional.
Actionable Practices
| Guardrail | How | When |
|---|---|---|
| Action clipping | Hard bounds on actions; fall back to safe values out of range | Design time |
| Shadow mode | New policy logs but doesn't act; compare for a week | Pre-launch |
| Gradual ramp-up | 1% -> 5% -> 50% staged traffic | Launch |
| Fallback rules | Keep a rule-based/manual takeover path on the business side | Always |
| Monitoring and alerts | Drift detection on action distributions, observation distributions, business metrics | Post-launch |
Safety design for RL can get very hardcore (constrained RL, safe sets), but the cheapest guardrail in engineering is always "don't let it touch the real system directly."
9. Principle 7: Sample-Efficiency Awareness
Counterexample: A team trained an on-policy algorithm (each interaction's sample is used exactly once) in a business environment that could only collect 50k samples. At sample 30k they realized the budget couldn't cover a single decent on-policy update cycle — they should have used an off-policy algorithm (SAC) or offline data.
In RL, "data" is generated by interaction, and every step has a real cost. Budget first, algorithm second.
Actionable Practices
- At project kickoff, answer three numbers: how many samples can be collected per hour, what's the total step budget, and what's the monetary cost per step.
- Pick the algorithm family by budget (the selection table in The Actor-Critic Family): a tight budget -> off-policy (SAC/TD3); a generous budget -> on-policy (PPO) is more stable.
- Report both dimensions — sample efficiency and final performance (see Building an RL Evaluation Suite from Scratch).
- When historical data exists, consider Offline RL first — squeeze whatever free samples you can.
10. Principle 8: Small-Step Iteration
Counterexample: On Monday an engineer changed the reward function, swapped the network, and tuned clip all at once. On Wednesday the results looked worse, with no way to tell which change did it — so everything was rolled back and the week was wasted.
RL experiments juggle too many variables; changing one thing at a time is the fastest path (the core discipline of Hyperparameter Tuning in Practice).
Actionable Practices
- Each experiment's change list has at most one variable (batch sweeps excepted — those are documented parallelism).
- One cycle a day: complete at least one loop of "change code -> run experiment -> inspect curves -> write a conclusion" every day.
- Leave a one-line written conclusion for every experimental round, even if it's just "lr 1e-3 -> collapsed."
The counterintuitive part of small steps
"Change one variable a day" looks slow, but it's actually the fastest way to move an RL project forward — attribution in RL is so expensive that with many variables, each experiment's information content approaches zero.
11. Principle 9: Document Failures
Counterexample: A team hit a "reward value overflow causing NaN" pitfall in Project A and spent three days diagnosing it. Six months later, a colleague on Project B independently stepped into the exact same pit.
Failure is reusable knowledge — and the most expensive kind. Fail to document it, and you pay full tuition for every failure.
Actionable Practices
- Maintain a
KNOWN_FAILURES.md: symptom -> root cause -> detection method -> fix, one entry per line. - Every time a bug takes more than a day to fix, add an entry.
- Link it from the README or wiki, and make it required reading for new members.
This is exactly what the whole Common Pitfalls and Anti-Patterns page does — treat it as the seed of your team's knowledge base and add your own entries to it.
12. Principle 10: Align with the Business
Counterexample: An inventory-replenishment project used "cumulative profit in the simulator" as its metric and optimized it to +18%. After launch, the warehouse manager reported: whenever predicted demand fluctuated, the policy's order quantities swung wildly, shift scheduling broke down, and real profit gained +0%. The metric had never been translated into the business constraint (order smoothness).
A win on RL metrics is not a business win. Between the two sit constraints, risks, and the cost of explanation.
Actionable Practices
- Metric translation table: write down, for every RL metric, what business action it corresponds to (return -> profit, success rate -> fulfillment rate, variance -> scheduling volatility).
- Constraints first: encode business constraints (smoothness, inventory caps, ethical boundaries) into the reward or action space up front — don't wait for the business side to discover them after launch.
- Explanation budget: the business will ask "why did it do that?" RL policies are notoriously hard to interpret, so prepare answers in the form of metric attribution or behavior audits in advance.
13. The Project Self-Check Checklist
Print the checklist below and pin it above your desk. Review it at the end of every phase:
At kickoff
- [ ] A decision checklist confirms this is an RL problem (step 0 in Build an RL Project from Scratch)
- [ ] The three sample-budget questions (steps / cost / hourly throughput) are answered
- [ ] The reward-design document is written, and the "score-farming audit" is done
- [ ] The business-metric translation table is set
During development
- [ ] All environment acceptance checks pass, and the heuristic baseline scores
- [ ] The training pipeline has been validated on a CartPole-style benchmark
- [ ] Baseline curves appear on every results plot
- [ ] Three-layer seeding and config archiving are in place
Before launch
- [ ] Multi-seed evaluation with an IQR report is complete
- [ ] Safety guardrails (clipping / shadow mode / fallback / monitoring) are in place
- [ ] The failure document is updated through the latest entry
Further Reading
- Reward Engineering — the full methodology behind Principle 1: shaping, sparse rewards, reward hacking, inverse RL
- Common Pitfalls and Anti-Patterns — the failure points corresponding to the ten principles and how to detect them
- LLM Alignment: RLHF in Practice — the extreme form of reward-design principles in LLM alignment (reward over-optimization)
- Anatomy of an RL System — the six-layer system architecture and the theoretical home of "the environment is the product"
- Build an RL Project from Scratch — the eight-step pipeline that puts these principles into practice
References
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. MIT Press. (the theoretical foundation of the reward hypothesis and trial-and-error learning; free version at incompleteideas.net)
- Amodei, D. et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565 (the classic list of safety-guardrail problems: reward misspecification, exploration risk, etc.)
- Irpan, A. (2018). Deep Reinforcement Learning Doesn't Work Yet. alexirpan.com/2018/02/14/rl-hard.html
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
- Tobin, J. et al. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. arXiv:1703.06907 (how to handle environment uncertainty — an extension of Principle 2)
- OpenAI (2017). Learning Dexterous In-Hand Manipulation. arXiv:1808.00177 (an industrial-grade case of environment-first + domain randomization)