Appearance
Building an RL Project from Scratch
In one sentence: this page walks through the complete pipeline of an RL project, from "an idea in your head" to "a policy running in production" — it answers the "I know the algorithms, but in what order do I actually build this?" question. It's written for engineers tackling their first real RL project; by the end, you'll be able to sketch your own project roadmap and sidestep the most common rework traps.
The defining feature of an RL project is its inverted cost structure. In a supervised learning project, data labeling is the most expensive part; in an RL project, the biggest cost by far is the environment and the reward, with the algorithm itself a distant second. Google's Anatomy of an RL System contains a much-quoted line: the environment is the product — roughly 80% of an RL project's cost, time, and risk of failure sits in environment and reward design, not in the PPO agent. Starting from that fact, this page breaks the project into eight sequential steps.
Project Workflow at a Glance: The Eight-Step Pipeline
┌────────────────────────────────────────────────────────────────┐
│ The Eight-Step RL Project Pipeline │
│ │
│ Step 0 Should you use RL at all? ← decision checklist; │
│ │ filters out ~30% of │
│ │ doomed projects upfront │
│ Step 1 Environment: pick / build / wrap ← highest-ROI step │
│ │ │
│ Step 2 Reward & observation design ← "the reward is the │
│ │ spec"; get it wrong │
│ │ and all is wasted │
│ Step 3 Baselines: random / heuristic / supervised │
│ │ ← no baselines, no algo │
│ Step 4 Minimal training loop ← make it run first, │
│ │ tune later │
│ Step 5 Evaluation & logging ← one protocol, used throughout │
│ │ │
│ Step 6 Reproducibility hardening ← seeds / config / deps │
│ │ │
│ Step 7 Deployment & monitoring ← serving, drift, rollback │
│ │ │
│ Iteration: Steps 4-6 form a loop you turn daily; mistakes │
│ in Steps 1-2 send you back upstream │
└────────────────────────────────────────────────────────────────┘These eight steps are not a waterfall. What you actually walk over and over is the inner loop formed by Steps 4-6 (train → evaluate → fix reproducibility issues, cycling several times a day), plus the occasional jump from that loop back to Steps 1-2 when you discover the environment or the reward is broken. Step 0 happens only once — so get it right.
The sections below follow the same order. This page cross-references Evaluation Protocols, Choosing a Framework, Tuning in Practice, and Common Pitfalls, among others — they're best read together.
Step 0: Should You Use RL at All?
RL is not a hammer. At project kickoff, run a decision checklist first and keep unsuitable projects out the door — that's where you save the most time and money. Here's a battle-tested one:
| Question | Yes | No | Notes |
|---|---|---|---|
| Can the problem be written as a state → action → reward loop? | Proceed | Pick another method | No sequential decision structure? Use supervised learning or rules |
| Is there a channel for repeated trial and error (simulator / sandbox / low-stakes production)? | Proceed | Pick another method | RL lives on trial and error; the cost of mistakes decides everything |
| Can the reward be defined cleanly? Can it be gamed? | Proceed | Do reward design first | A misspecified reward collapses harder the stronger the algorithm gets — see Reward Engineering |
| Do you have a budget of 10^4+ interaction samples? | Proceed | Consider offline RL / BC | Too few samples to even beat a baseline |
| Does a simple heuristic or rule already come close to the goal? | Probably don't use RL | Proceed | If you can write a rule, write the rule |
| Can the business tolerate the policy being bad for a while before it gets good? | Proceed | Pick another method | RL performance swings wildly early on — see RL Design Principles |
The single most common mistake
Treating "RL could work here" as "RL should be used here." A real case: a team built an "intelligent scheduler" for warehouse order picking. The requirement was "5% faster than the manual rules," but every real order had to be replayed in simulation, and a single evaluation took 40 minutes — sample efficiency could never support training. In the end, a simple sequencing heuristic hit the target. Ask "what's the cheapest solution?" before asking "can RL do better?"
Three verdicts that rule RL out:
- No sequential structure: each decision is independent and settles things in one shot → plain classification/regression or a multi-armed bandit will do (Multi-Armed Bandits are a degenerate form of RL).
- No trial-and-error channel: consequences are irreversible and there's no simulator → use behavior cloning or Offline RL on logged data; never online RL.
- Exactly modelable: you can write down the dynamics and the problem is small → dynamic programming or optimal control (RL vs. Neighboring Fields).
When RL is clearly the right answer
Cheap simulation (games, robot simulators), a reward that comes for free (score, win rate, profit), a clearly defined action space, and a solution space too large to cover with hand-written rules — meet two or more of these, and RL is usually worth a shot. See the project suggestions in Learning Paths.
The Environment: Choosing, Building, Wrapping
The environment is the single biggest source of risk in an RL project. It determines how many samples a day of training buys you, whether your evaluation results are trustworthy, and whether the learned policy transfers at all.
1. Three Sources
| Source | Best for | Cost | Risk |
|---|---|---|---|
| Off-the-shelf environment library | The project is fundamentally about "running algorithms" (learning, reproduction, comparison) | Low | Few environment bugs, but the env may not match your business |
| Off-the-shelf + customization | Your business closely resembles a public environment (e.g., scheduling ≈ job-shop) | Medium | The custom parts introduce new bugs |
| Fully custom | Unique to your business (recommendation, trading, physical systems) | High | The environment is the product — every bug hides here |
For choosing among environment libraries, see the Datasets & Tools Index: Gymnasium is the de facto standard API, MuJoCo/Brax cover continuous control (and Brax runs on GPU), while Isaac-style platforms handle robotics. Whichever you pick, the very first task is a 30-line "smoke test" for the environment: run 1,000 steps with random actions and verify that step's return shapes, the conditions that trigger done, and the reward range all match expectations.
2. Three Disciplines for Building Your Own Environment
A custom environment (a recommendation simulator, a scheduling simulation, and so on) is exactly what "the environment is the product" refers to. Three disciplines are non-negotiable:
- Define the interface before writing the logic. Subclass
gymnasium.Envdirectly and implementreset,step,observation_space, andaction_space. For Gymnasium API details, see the Progressive Tutorial. - Declare observations and actions explicitly with
spacesand validate withspace.contains(). Getting the declaration wrong (e.g., a discrete action written as continuous) is a top cause of training failing silently. - Keep the environment separate from the training code. The environment is a standalone process or service, and the training code only talks to it through the interface. That way you can swap implementations (Python → C++ → distributed) without ever touching the algorithm code.
3. A Minimal Custom Environment
Below is a toy "1D cart moves toward a goal" environment that shows the interface skeleton:
python
import gymnasium as gym
import numpy as np
from gymnasium import spaces
class ReachGoalEnv(gym.Env):
"""1D cart motion: the state is position and velocity, the action is a
force, and the goal is to come to rest near [0, 0]."""
metadata = {"render_modes": []}
def __init__(self):
super().__init__()
# Action: a continuous force in [-1, 1]
self.action_space = spaces.Box(low=-1.0, high=1.0, shape=(1,), dtype=np.float32)
# Observation: position and velocity
self.observation_space = spaces.Box(low=-np.inf, high=np.inf, shape=(2,), dtype=np.float32)
self._state = np.zeros(2, dtype=np.float32)
self._step_count = 0
def reset(self, *, seed=None, options=None):
super().reset(seed=seed) # required: seeds np_random
self._state = np.array([self.np_random.uniform(-1, 1), 0.0], dtype=np.float32)
self._step_count = 0
return self._state.copy(), {}
def step(self, action):
action = np.clip(action, -1.0, 1.0).item()
pos, vel = self._state
vel = vel + 0.1 * action # simple dynamics
pos = pos + vel
self._state = np.array([pos, vel], dtype=np.float32)
self._step_count += 1
# Reward: closer to the goal is better; also encourages coming to rest
reward = -abs(pos) - 0.5 * abs(vel)
# Termination: near the goal or out of steps
terminated = abs(pos) < 0.05 and abs(vel) < 0.05
truncated = self._step_count >= 200
info = {"dist": abs(pos)}
return self._state.copy(), reward, terminated, truncated, infoMemorize the interface conventions
In current Gymnasium, reset returns (obs, info) and step returns (obs, reward, terminated, truncated, info), and terminated (task completed) must be kept separate from truncated (cut short by external forces, such as a time limit). Treating a timeout as "task failed" or "task succeeded" systematically distorts the learning signal. See the tutorial's pitfall list.
4. Environment Acceptance Checklist
Before a custom environment is signed off, walk through every item:
- [ ] 1,000 rollouts under a random policy raise no exceptions; observation/reward types match the space declarations
- [ ] Rewards are bounded (or at least have bounded variance), so reward normalization can be designed later
- [ ]
resetproduces the same initial-state sequence for the same seed - [ ] A simple hand-written heuristic scores points, confirming the environment is solvable (i.e., the environment itself isn't impossible)
- [ ]
stepthroughput has been measured (steps per second) to confirm it supports the target sample budget
Steps 2 and 3: Reward Design and Baselines First
1. Design the Reward Before Picking the Algorithm
"The reward is the spec" — the reward function is the only machine-readable expression of the product requirements. Get this step wrong and everything downstream is wasted; see Reward Engineering for the full treatment. At the practice level there's exactly one rule: write a reward design document first — for every reward term, list its source, its value, why it's set that way, and how it might be gamed. Only then start writing code.
2. Baselines First: No Baselines, No Algorithm
Many projects fire up SAC or PPO on day one, the agent doesn't learn, and nobody can tell whether the algorithm or the environment is at fault. The right move is to set up a ladder of difficulty:
| Level | Policy | Purpose | Expected outcome |
|---|---|---|---|
| 0 | Random policy | Verify the environment responds and rewards flow | Establishes the "floor" |
| 1 | Hand-written heuristic (rules / greedy / linear) | Verify the problem is solvable at all | Establishes a "reference line" |
| 2 | Supervised warm-up (behavior cloning / supervised baseline) | Reach a reasonable level quickly when offline data exists | A "cheap upper bound" |
| 3 | Simple RL (DQN / REINFORCE) | Verify the RL signal path works | At least matches the heuristic |
| 4 | Full RL (PPO / SAC + tuning) | Push toward the ceiling | Significantly beats the heuristic |
Why run the heuristic first
A heuristic is a ruler: it tells you directly whether the environment's reward signal is dense enough to learn from. If the heuristic scores 80 with ease while RL sits at 30, the problem almost certainly lies in the reward or observation design — not the algorithm.
An all-too-common mistake is skipping Level 1, going straight to PPO, and spending two weeks tuning hyperparameters before discovering the hand-written rule was better all along. Time saved at Level 1 always gets repaid with interest in Steps 4-6.
The Minimal Training Loop
With the environment and baselines in place, it's time to build the training loop. The principle is minimal skeleton, then incremental polish: the first version only needs to do three things — run, save the model, and produce a learning curve. Every optimization (vectorization, parallelism, distributed training) waits until the signal path is confirmed correct.
1. The Minimal Skeleton: A Framework-Agnostic Loop
python
"""Minimal RL training-loop skeleton: depends only on gymnasium + numpy.
Goal: run end to end, save the model, produce a curve. Swapping algorithms
does not change the skeleton's structure."""
import gymnasium as gym
import numpy as np
class RandomPolicy:
"""Placeholder policy: swap in any algorithm exposing .act(obs)."""
def __init__(self, env):
self.env = env
def act(self, obs):
return self.env.action_space.sample()
def collect_episode(env, policy, max_steps=500):
"""Run one episode, return (total return, number of steps)."""
obs, _ = env.reset()
total_reward, steps = 0.0, 0
for _ in range(max_steps):
action = policy.act(obs)
obs, reward, terminated, truncated, _ = env.step(action)
total_reward += reward
steps += 1
if terminated or truncated:
break
return total_reward, steps
def train(env_id, num_episodes=2000, seed=0):
env = gym.make(env_id)
# 1. Seed the environment (seed every environment)
env.reset(seed=seed)
policy = RandomPolicy(env)
returns = []
for ep in range(num_episodes):
# 2. Re-seed the environment every episode for reproducibility
env.reset(seed=seed + ep)
total_reward, _ = collect_episode(env, policy)
returns.append(total_reward)
# 3. Logging: print the window mean every 100 episodes (coarse first)
if (ep + 1) % 100 == 0:
mean = float(np.mean(returns[-100:]))
print(f"episode {ep+1:>5d} | mean_return(100) = {mean:.2f}")
# 4. Save the model + the curve
np.save("returns.npy", np.array(returns))
print("done, returns saved to returns.npy")
env.close()
if __name__ == "__main__":
train("CartPole-v1")Three design decisions in this skeleton are worth remembering:
- The policy object exposes a single interface,
act(obs)— DQN, REINFORCE, and PPO can all plug in the same way, so swapping algorithms barely touches the training loop. For the algorithm internals, see Value-Based Learning and Policy Gradient Methods. - Every episode re-seeds
reset— this separates environment randomness from algorithm randomness, so when something goes wrong you know which layer to blame. - Logging starts coarse and gets finer — the first version prints only a windowed mean; once the signal checks out, add diagnostics like mean TD error, entropy, and KL gradually (see Tuning in Practice).
2. From Skeleton to Production: Three Stages
| Stage | What you do | When |
|---|---|---|
| Skeleton | Single process, step-by-step episodes; confirm learning happens | Day one |
| Vectorized | Parallel collection with gymnasium.vector or SubprocVecEnv | When single process is too slow |
| Framework | Rewrite with, or port to, SB3 / CleanRL / RLlib | Once the skeleton proves it works |
Don't rush to a framework
Write a 50-line training loop by hand first, then go read the framework comparison. The most common outcome of reaching for a framework immediately is that the algorithm runs, but you don't understand a single number in the logs, and when a bug hits you have no idea where to start.
Evaluation and Logging
Decide the evaluation protocol at the start of the project, not after training — because every "this looks better" decision you make during training rests on an evaluation judgment. The full protocol lives in Evaluation in Practice; here are the three minimal pieces that actually work:
- A fixed evaluation function: every N environment steps during training, run
num_eval_episodesepisodes with fixed seeds and log "mean eval return ± interquartile range," kept separate from training returns. - Persist the learning curves: append
(timestep, train_return, eval_return, seed)to a CSV on every run — this is the raw material for all later diagnosis. - Write an experiment record for every run: config (hyperparameters, seed, environment), git commit, log path. An unrecorded experiment is an experiment that didn't happen.
The core concepts behind evaluation — sample efficiency vs. final performance, multi-seed matrices, IQR — come from Evaluation and Benchmarks; the practice-level template is in Building an RL Evaluation Suite from Scratch.
Hardening Reproducibility
Reproducibility in RL is an order of magnitude worse than in supervised learning: identical code and hyperparameters can produce different curves on a different machine, a different GPU, or even a different NumPy version. Here are the countermeasures, ranked by ROI:
1. Three-Layer Seed Management
| Layer | What you seed | Code |
|---|---|---|
| Global | random, numpy, torch | random.seed(seed); np.random.seed(seed); torch.manual_seed(seed) |
| Environment | Gymnasium environments | env.reset(seed=s); env.action_space.seed(s) |
| Inside the algorithm | Replay sampling, noise, action sampling | The framework's built-in seed(s) (e.g., SB3's PPO(..., seed=...)) |
A global seed is not a silver bullet
Some PyTorch operators are not deterministic on GPU (anything built on atomicAdd-style operations), and multi-GPU or parallel sampling makes bit-level reproduction flatly impossible. The pragmatic goal isn't "reproduce bit for bit" but reproduce the performance distribution for a given config (same mean and IQR) — and that requires multiple seeds, not one. See the seed pitfalls in Tuning in Practice.
2. Configuration Management: Hydra
Hyperparameters scattered through the code are the biggest hygiene problem in RL projects. Use a configuration system to separate "experiments" from "code":
yaml
# config/train.yaml (Hydra config example)
seed: 42
env_id: CartPole-v1
algorithm:
name: ppo
learning_rate: 3.0e-4
gamma: 0.99
gae_lambda: 0.95
clip_range: 0.2
ent_coef: 0.0
n_steps: 2048
batch_size: 64
evaluation:
eval_episodes: 10
eval_freq: 10000 # evaluate every N environment steps
seeds: [0, 1, 2] # multiple evaluation seedsOverride any field from the command line (python train.py algorithm.learning_rate=1e-3), and archive the full config of every experiment alongside its logs. "Reproducing a result" then degenerates into "re-running a config."
3. Locking Dependencies and Versions
- Pin
requirements.txt/pyproject.tomlto minor versions (RL code couples tightly to specific NumPy/Gymnasium versions). - Record the versions of
gymnasium,numpy,torch, andstable-baselines3in every experiment record. - If you can, freeze a Docker image with
docker build, or use the lockfiles ofuv/poetry.
Deployment and Monitoring
Training a good policy is only half the project. The classic failure mode at launch is: great offline scores, terrible online performance. The usual root causes are distribution shift and a mismatch between evaluation and deployment.
1. Serving Patterns for Policies
| Scenario | Serving pattern | Examples |
|---|---|---|
| Offline batch decisions | Batch job run on a schedule | Production scheduling, batch recommendation |
| Online single decisions | Low-latency inference service | Ad bidding, real-time control |
| Online environment interaction | The environment lives in the business system; the policy is called back | Robotics, dialogue systems |
The universal practice: export the policy to an inference-only format (ONNX / TorchScript / JAX), so the serving stack depends only on the "observation → action" forward function and carries no training dependencies whatsoever. The training loop, the replay buffer, and the optimizer have no business being in the serving path.
2. The Monitoring Trinity
- Action distribution: drift in the mean/variance/spread of online actions means the observation distribution has shifted.
- Observation distribution: compute statistics on live observations (means, quantiles, approximate hashes) and compare them against the training distribution.
- Business metrics: rewards often can't be measured online directly (delayed rewards, human judgment), so proxy business metrics must be watched at the same time.
text
Drift detection flow (weekly patrol):
collect this week's live observations → compare distributions against
the training data (KS test / histograms)
→ significant difference? → yes: trigger a retrain evaluation;
no: keep watching3. Rollback and Safety Rails
- Version every policy, and keep the last N versions in the model server so you can roll back with one click.
- Before going live, run a shadow deployment: the policy only logs its actions without executing them, and you compare against the existing system for a week.
- Keep a rule-based fallback on the business side: when a policy action is abnormal (out of range, NaN, timeout), fall back to the rule policy. Safety-related discussion is in RL Design Principles and Common Pitfalls.
Iteration Rhythm: Make It Run First, Then Tune
To close, here's an executable rhythm sheet. Most failed RL projects share one trait — not "the algorithm didn't work," but tuning before the pipeline actually ran, or tuning without discipline after it did.
| Phase | Suggested duration | Deliverable | Exit condition |
|---|---|---|---|
| Steps 0-2 | 1-2 weeks | Decision checklist, environment acceptance sign-off, reward design doc | Environment runs; heuristic scores |
| Steps 3-4 | 3-5 days | Skeleton training loop, random/heuristic baseline curves | Can save models and produce learning curves |
| Steps 5-6 | 3-5 days | Evaluation script, CSV logs, Hydra configs | One command fully reproduces an experiment |
| Step 7 | In parallel with training | Deployment pipeline, monitoring and alerts | Shadow deployment works |
| Tuning loop | Weekly | Change one variable per round; keep experiment records | Business acceptance criteria met |
The one-loop-a-day discipline
Treat "change code → run experiment → read curves → write down conclusions" as a single loop and complete at least one loop per day. RL experiments routinely take hours, and the nightmare scenario is changing five variables, running for five hours, and ending up unable to say which variable did anything. Single-variable experiments, make-it-run-before-tuning, and experiment records — these three habits alone will save you 60% of your total project time.
Further Reading
- Anatomy of an RL System — the six-layer architecture of a complete RL system; the theoretical foundation of this page's pipeline
- Progressive Tutorial: Three Gymnasium Versions from Scratch — the full flesh on this page's skeleton; run all three versions yourself
- Building an RL Evaluation Suite from Scratch — the full expansion of Step 6 on this page: experiment matrices, IQR, fairness traps
- How to Choose Frameworks and Tools — once the skeleton matures, how to graduate to SB3 / RLlib / CleanRL
- Common Pitfalls and Anti-Patterns — the failure modes and detection methods for each step on this page
- Datasets & Tools Index — a one-stop index of environment libraries, benchmark suites, and training toolchains
References
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. MIT Press. (The whole book, especially Chapters 1, 2, and 8 on the agent-environment interface and tabular methods; free online edition at incompleteideas.net/book/the-book-2nd.html)
- Gymnasium official documentation (Farama Foundation): gymnasium.farama.org;
Envinterface reference at gymnasium.farama.org/api/env/ - Stable-Baselines3 official documentation: stable-baselines3.readthedocs.io
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560 (the classic treatment of RL reproducibility)
- Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133 (empirical evidence on seed sensitivity)
- Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292 (why implementation details matter more than the paper's equations)
- Irpan, A. (2018). Deep Reinforcement Learning Doesn't Work Yet. alexirpan.com/2018/02/14/rl-hard.html (the famous blog post on why RL projects fail)
- Hydra official documentation: hydra.cc