Skip to content

RL Design Principles

On this page Ten principles grown out of failed projects — reward before algorithm, environment first, simplicity first, baselines first, reproducibility, safety guardrails, sample-efficiency awareness, iteration cadence, documenting failures, business alignment.

RL Design Principles ​

One-liner: this page distills the lessons of repeatedly failed RL projects into ten design principles — each paired with a real-world counterexample and one actionable practice. It addresses the root cause of "why is my RL project always stuck in overtime rework," and it serves tech leads, solo developers, and anyone who has to actually ship an RL system. By the end, you'll walk away with two pieces of paper: the ten principles to pin above your desk, and a project self-check checklist.

Start with the facts. Why do RL projects fail significantly more often than ordinary ML projects? Because RL failure modes are more numerous, better hidden, and surface later:

  • Failures hide better: when training stalls, you can't tell whether it's the algorithm, the hyperparameters, the reward, or an environment bug (Common Pitfalls covers this in full).
  • Failures arrive late: supervised learning reveals bad data within hours; RL often runs for a day or two before you discover it "learned the wrong signal."
  • Failures cost more: trial and error needs an environment, and the more realistic the environment, the pricier — robots, trading, and real users all carry enormous trial-and-error costs.

These ten principles grew out of exactly those failures. They aren't inspiration — they're a memo against paying the same tuition twice.

1. Why RL Projects Fail: Three "Failure Bills" ​

Failure typeTypical share (informal estimate)Representative symptomsRelated principles
Wrong problem formulation~30%"This isn't an RL problem at all," or a mis-specified rewardPrinciples 1, 2
Engineering problems~30%Environment bugs, irreproducible results, lost logsPrinciples 2, 5, 7
Algorithm / tuning problems~25%Won't learn, collapses, high variancePrinciples 3, 4, 6
Business / deployment problems~15%Built but unused; metrics don't match the businessPrinciples 8, 9, 10

Note: only a quarter of the problems are "algorithm" problems. Yet the vast majority of team resources get poured into algorithms — which is precisely why Principle 1 exists.

2. The Ten Principles at a Glance ​

#PrincipleIn one lineThe counterexample
1Reward before algorithmThe reward is the product spec — write the reward doc before picking an algorithmA mis-specified reward wastes even the best PPO
2Environment firstThe environment is the product; its quality decides everythingOne environment bug wastes two weeks for everyone
3Simplicity firstStart with the simplest viable strategy and upgrade step by stepReach for SAC on day one; three days of silence
4Baselines firstWithout a baseline, "getting better" is unverifiableNo random/heuristic baseline -> a false "the algorithm works"
5Reproducible experimentsAn experiment you can't reproduce is one you didn't runResults flip completely on a different machine
6Safety guardrailsDecide what failure looks like before going liveThe policy flails and production goes down
7Sample-efficiency awarenessEvery interaction has a cost — budget firstA 10k-step training budget with an on-policy algorithm
8Small-step iterationChange one thing at a time, one cycle a dayFive variables changed at once; two hours later, no attribution
9Document failuresTurn every failure into a lessons libraryThe same pitfall gets stepped into by three separate teams
10Align with the businessMetrics must translate into business languageGreat algorithm scores, unmoved business metrics

3. Principle 1: Design the Reward Before Choosing the Algorithm ​

Counterexample: A team designed a reward to "stop NPCs from getting stuck on walls" — "+1 for every step closer to the goal, -10 for hitting a wall." The agent learned to hug the wall (+1 every step is a sure thing; the wall-collision probability is low). The better the algorithm optimized, the more "cleverly bad" the behavior became.

The reward is the only machine-readable expression of the product requirements. Your choice of algorithm only determines whether learning happens at all; the reward determines what the learned thing looks like.

Actionable Practices ​

  1. Before writing any code, draft a one-page reward-design document: the source, value, and rationale of each reward term, plus the loopholes it might invite.
  2. Health-check the reward magnitude: run a random policy and look at the average per-step reward — if it's on the order of 1e-4, floating-point precision alone is polluting learning; normalize first.
  3. Run a "reward audit" before launch: enumerate the cheapest score-farming paths the policy might find, and close them one by one.

The full methodology of reward design (potential-based shaping, sparse-reward countermeasures, a case collection of reward hacking, inverse RL) lives in Reward Engineering. In RLHF settings the "reward" becomes a model of human preferences, and this principle applies just as much — over-optimizing the reward model is precisely an alignment failure; see LLM Alignment: RLHF in Practice.

The one-sentence test

Explain your reward function to a non-technical person. If they can point out "how the agent could cheat for points," your reward still has holes.

4. Principle 2: Environment First (The Environment Is the Product) ​

Counterexample: A scheduling project trained in a home-grown simulator for four weeks; the agent performed flawlessly. At launch they discovered the simulator's "machine failure probability" was hard-coded at 5%, while the real production line ran at 15% and drifted over time — the policy collapsed on the true distribution.

The line "the environment is the product" in Anatomy of an RL System is not a slogan: the environment determines your data distribution, and the data distribution determines what the policy can learn. Get the environment wrong, and everything after is wasted effort.

Actionable Practices ​

  1. Environment acceptance checklist (the full list is in Build an RL Project from Scratch): smoke tests, a random policy running 1,000 steps, bounded rewards, same-seed reproducibility, a reachable heuristic.
  2. Verify "the environment is trustworthy" before writing any algorithm: run the simplest possible policy to confirm a learning signal exists before spending algorithm effort.
  3. Decouple the environment from training: keep it as a standalone service/module so swapping implementations (Python -> C++ -> distributed) never touches training code.
  4. Domain-randomization awareness: the simulator and reality have a gap, so randomize physics parameters during training (robotics scenarios: see Robot Control and Sim2Real).

The environment is also the requirements spec

Write the environment like a product requirements document: what the states are, action bounds, the reward formula, termination conditions, sources of randomness. That document doubles as test cases, a review artifact, and handoff material — it deserves more review rounds than the algorithm code.

5. Principle 3: Start Simple (Random -> Heuristic -> Simple RL -> Complex RL) ​

Counterexample: A team reached for SAC the moment the project was approved, with 8 GPUs running parallel tuning. Two weeks later, the return hadn't budged. They switched to the plainest ε-greedy linear policy, which caught up with the heuristic baseline in a single day — the problem had never needed SAC's complexity in the first place.

Complexity is a liability: the more complex the algorithm, the harder to debug, reproduce, and explain. Run the simplest thing first, until simplicity is proven insufficient.

The Complexity Ladder ​

text
Random policy ──→ Heuristic/rules ──→ Simple RL (DQN/REINFORCE)
      │                  │                     │
 Verify the env      Verify the problem   Verify the RL signal path
   is interactive     is solvable                 │
                                            Complex RL (PPO/SAC) + tuning
                                                    │
                                            Upgrade only when simple falls short
RungPurposeWhen to move down
RandomEnvironment smoke test, reward collectionImmediately
HeuristicProblem solvability, reward sanityOnce it scores
Simple RLRL signal path, correct training pipelineOnce the pipeline is verified
Complex RLPush the ceilingOnce simple is proven insufficient

6. Principle 4: Baselines First ​

Counterexample: The team tuned PPO for over a month; the final return was 3x the random policy, and everyone celebrated. Then they ran the heuristic baseline: a 50-line greedy rule achieved 4.5x random. Three months of "algorithm work" had added nothing.

"Baselines first" is the cheapest, most life-saving principle: it turns "is my algorithm any good?" into an answerable question.

Actionable Practices ​

  1. Within the first week, produce complete learning curves for a random baseline and a heuristic baseline.
  2. Overlay every subsequent experiment on these two lines (two plt.axhline reference lines are enough).
  3. Keep the evaluation protocol identical to the baselines' (same seeds, budget, metrics) — otherwise the baseline is meaningless. Protocol details: Building an RL Evaluation Suite from Scratch.

Don't make baselines a formality

If "baselines first" happens once and is never updated, it never happened. The baseline's value lies in being a standing reference: every algorithm change must answer "where is it better than the baseline, and by how much?" Keep the baseline lines on every plot until the project ends.

7. Principle 5: Reproducible Experiments ​

Counterexample: Engineer A ran PPO with PyTorch 1.13 and got 480 points; Engineer B reproduced the same code in a fresh environment with PyTorch 2.0 and got 390. After three days of arguing, they found the cause: the numpy random state had never been seeded inside the environment — the same seed followed different environment sequences on the two machines.

Reproducibility in RL is worse than in supervised learning: same seed, same code — change the GPU, the library version, or the parallelism scheme, and the results change.

Actionable Practices ​

  1. Three-layer seeding: global (python/numpy/torch) + environment (reset(seed=...)) + algorithm (framework seed parameter).
  2. Config as code: experiment configs live in YAML/Hydra, can be overridden from the command line, and are archived automatically with every run.
  3. Record everything: git commit, dependency versions, and machine info go into the experiment directory.
  4. Multiple seeds are part of reproducibility: what you reproduce is the "performance distribution," not "one curve."

Full engineering details are in the reproducibility-hardening section of Build an RL Project from Scratch; the diagnostic checklist for "my paper reproduction failed" is in Common Pitfalls.

8. Principle 6: Safety Guardrails ​

Counterexample: A ranking policy for recommendations looked good in offline evaluation only, with no shadow mode. On its first live day, the policy took extreme actions for a small slice of users (recommending the same product over and over), CTR cratered, and the rollback took three hours.

RL policies are more likely than rule-based systems to "cleverly cross the line." Guardrails are not optional.

Actionable Practices ​

GuardrailHowWhen
Action clippingHard bounds on actions; fall back to safe values out of rangeDesign time
Shadow modeNew policy logs but doesn't act; compare for a weekPre-launch
Gradual ramp-up1% -> 5% -> 50% staged trafficLaunch
Fallback rulesKeep a rule-based/manual takeover path on the business sideAlways
Monitoring and alertsDrift detection on action distributions, observation distributions, business metricsPost-launch

Safety design for RL can get very hardcore (constrained RL, safe sets), but the cheapest guardrail in engineering is always "don't let it touch the real system directly."

9. Principle 7: Sample-Efficiency Awareness ​

Counterexample: A team trained an on-policy algorithm (each interaction's sample is used exactly once) in a business environment that could only collect 50k samples. At sample 30k they realized the budget couldn't cover a single decent on-policy update cycle — they should have used an off-policy algorithm (SAC) or offline data.

In RL, "data" is generated by interaction, and every step has a real cost. Budget first, algorithm second.

Actionable Practices ​

  1. At project kickoff, answer three numbers: how many samples can be collected per hour, what's the total step budget, and what's the monetary cost per step.
  2. Pick the algorithm family by budget (the selection table in The Actor-Critic Family): a tight budget -> off-policy (SAC/TD3); a generous budget -> on-policy (PPO) is more stable.
  3. Report both dimensions — sample efficiency and final performance (see Building an RL Evaluation Suite from Scratch).
  4. When historical data exists, consider Offline RL first — squeeze whatever free samples you can.

10. Principle 8: Small-Step Iteration ​

Counterexample: On Monday an engineer changed the reward function, swapped the network, and tuned clip all at once. On Wednesday the results looked worse, with no way to tell which change did it — so everything was rolled back and the week was wasted.

RL experiments juggle too many variables; changing one thing at a time is the fastest path (the core discipline of Hyperparameter Tuning in Practice).

Actionable Practices ​

  • Each experiment's change list has at most one variable (batch sweeps excepted — those are documented parallelism).
  • One cycle a day: complete at least one loop of "change code -> run experiment -> inspect curves -> write a conclusion" every day.
  • Leave a one-line written conclusion for every experimental round, even if it's just "lr 1e-3 -> collapsed."

The counterintuitive part of small steps

"Change one variable a day" looks slow, but it's actually the fastest way to move an RL project forward — attribution in RL is so expensive that with many variables, each experiment's information content approaches zero.

11. Principle 9: Document Failures ​

Counterexample: A team hit a "reward value overflow causing NaN" pitfall in Project A and spent three days diagnosing it. Six months later, a colleague on Project B independently stepped into the exact same pit.

Failure is reusable knowledge — and the most expensive kind. Fail to document it, and you pay full tuition for every failure.

Actionable Practices ​

  1. Maintain a KNOWN_FAILURES.md: symptom -> root cause -> detection method -> fix, one entry per line.
  2. Every time a bug takes more than a day to fix, add an entry.
  3. Link it from the README or wiki, and make it required reading for new members.

This is exactly what the whole Common Pitfalls and Anti-Patterns page does — treat it as the seed of your team's knowledge base and add your own entries to it.

12. Principle 10: Align with the Business ​

Counterexample: An inventory-replenishment project used "cumulative profit in the simulator" as its metric and optimized it to +18%. After launch, the warehouse manager reported: whenever predicted demand fluctuated, the policy's order quantities swung wildly, shift scheduling broke down, and real profit gained +0%. The metric had never been translated into the business constraint (order smoothness).

A win on RL metrics is not a business win. Between the two sit constraints, risks, and the cost of explanation.

Actionable Practices ​

  1. Metric translation table: write down, for every RL metric, what business action it corresponds to (return -> profit, success rate -> fulfillment rate, variance -> scheduling volatility).
  2. Constraints first: encode business constraints (smoothness, inventory caps, ethical boundaries) into the reward or action space up front — don't wait for the business side to discover them after launch.
  3. Explanation budget: the business will ask "why did it do that?" RL policies are notoriously hard to interpret, so prepare answers in the form of metric attribution or behavior audits in advance.

13. The Project Self-Check Checklist ​

Print the checklist below and pin it above your desk. Review it at the end of every phase:

At kickoff

  • [ ] A decision checklist confirms this is an RL problem (step 0 in Build an RL Project from Scratch)
  • [ ] The three sample-budget questions (steps / cost / hourly throughput) are answered
  • [ ] The reward-design document is written, and the "score-farming audit" is done
  • [ ] The business-metric translation table is set

During development

  • [ ] All environment acceptance checks pass, and the heuristic baseline scores
  • [ ] The training pipeline has been validated on a CartPole-style benchmark
  • [ ] Baseline curves appear on every results plot
  • [ ] Three-layer seeding and config archiving are in place

Before launch

  • [ ] Multi-seed evaluation with an IQR report is complete
  • [ ] Safety guardrails (clipping / shadow mode / fallback / monitoring) are in place
  • [ ] The failure document is updated through the latest entry

Further Reading ​

References ​

  • Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. MIT Press. (the theoretical foundation of the reward hypothesis and trial-and-error learning; free version at incompleteideas.net)
  • Amodei, D. et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565 (the classic list of safety-guardrail problems: reward misspecification, exploration risk, etc.)
  • Irpan, A. (2018). Deep Reinforcement Learning Doesn't Work Yet. alexirpan.com/2018/02/14/rl-hard.html
  • Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
  • Tobin, J. et al. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. arXiv:1703.06907 (how to handle environment uncertainty — an extension of Principle 2)
  • OpenAI (2017). Learning Dexterous In-Hand Manipulation. arXiv:1808.00177 (an industrial-grade case of environment-first + domain randomization)