Skip to content

Building an RL Project from Scratch

On this page The complete pipeline of an RL project, from problem definition to production — choosing an environment (build vs. off-the-shelf), baselines first, the training loop, evaluation protocols, logging and reproducibility, deployment and monitoring — at the "make it run first, tune later" rhythm.

Building an RL Project from Scratch ​

In one sentence: this page walks through the complete pipeline of an RL project, from "an idea in your head" to "a policy running in production" — it answers the "I know the algorithms, but in what order do I actually build this?" question. It's written for engineers tackling their first real RL project; by the end, you'll be able to sketch your own project roadmap and sidestep the most common rework traps.

The defining feature of an RL project is its inverted cost structure. In a supervised learning project, data labeling is the most expensive part; in an RL project, the biggest cost by far is the environment and the reward, with the algorithm itself a distant second. Google's Anatomy of an RL System contains a much-quoted line: the environment is the product — roughly 80% of an RL project's cost, time, and risk of failure sits in environment and reward design, not in the PPO agent. Starting from that fact, this page breaks the project into eight sequential steps.

Project Workflow at a Glance: The Eight-Step Pipeline ​

┌────────────────────────────────────────────────────────────────┐
│  The Eight-Step RL Project Pipeline                            │
│                                                                │
│  Step 0  Should you use RL at all?  ← decision checklist;      │
│     │                                 filters out ~30% of      │
│     │                                 doomed projects upfront  │
│  Step 1  Environment: pick / build / wrap  ← highest-ROI step  │
│     │                                                          │
│  Step 2  Reward & observation design  ← "the reward is the     │
│     │                                   spec"; get it wrong    │
│     │                                   and all is wasted      │
│  Step 3  Baselines: random / heuristic / supervised            │
│     │                                 ← no baselines, no algo  │
│  Step 4  Minimal training loop  ← make it run first,           │
│     │                             tune later                   │
│  Step 5  Evaluation & logging  ← one protocol, used throughout │
│     │                                                          │
│  Step 6  Reproducibility hardening  ← seeds / config / deps    │
│     │                                                          │
│  Step 7  Deployment & monitoring  ← serving, drift, rollback   │
│     │                                                          │
│  Iteration: Steps 4-6 form a loop you turn daily; mistakes     │
│  in Steps 1-2 send you back upstream                           │
└────────────────────────────────────────────────────────────────┘

These eight steps are not a waterfall. What you actually walk over and over is the inner loop formed by Steps 4-6 (train → evaluate → fix reproducibility issues, cycling several times a day), plus the occasional jump from that loop back to Steps 1-2 when you discover the environment or the reward is broken. Step 0 happens only once — so get it right.

The sections below follow the same order. This page cross-references Evaluation Protocols, Choosing a Framework, Tuning in Practice, and Common Pitfalls, among others — they're best read together.

Step 0: Should You Use RL at All? ​

RL is not a hammer. At project kickoff, run a decision checklist first and keep unsuitable projects out the door — that's where you save the most time and money. Here's a battle-tested one:

QuestionYesNoNotes
Can the problem be written as a state → action → reward loop?ProceedPick another methodNo sequential decision structure? Use supervised learning or rules
Is there a channel for repeated trial and error (simulator / sandbox / low-stakes production)?ProceedPick another methodRL lives on trial and error; the cost of mistakes decides everything
Can the reward be defined cleanly? Can it be gamed?ProceedDo reward design firstA misspecified reward collapses harder the stronger the algorithm gets — see Reward Engineering
Do you have a budget of 10^4+ interaction samples?ProceedConsider offline RL / BCToo few samples to even beat a baseline
Does a simple heuristic or rule already come close to the goal?Probably don't use RLProceedIf you can write a rule, write the rule
Can the business tolerate the policy being bad for a while before it gets good?ProceedPick another methodRL performance swings wildly early on — see RL Design Principles

The single most common mistake

Treating "RL could work here" as "RL should be used here." A real case: a team built an "intelligent scheduler" for warehouse order picking. The requirement was "5% faster than the manual rules," but every real order had to be replayed in simulation, and a single evaluation took 40 minutes — sample efficiency could never support training. In the end, a simple sequencing heuristic hit the target. Ask "what's the cheapest solution?" before asking "can RL do better?"

Three verdicts that rule RL out:

  1. No sequential structure: each decision is independent and settles things in one shot → plain classification/regression or a multi-armed bandit will do (Multi-Armed Bandits are a degenerate form of RL).
  2. No trial-and-error channel: consequences are irreversible and there's no simulator → use behavior cloning or Offline RL on logged data; never online RL.
  3. Exactly modelable: you can write down the dynamics and the problem is small → dynamic programming or optimal control (RL vs. Neighboring Fields).

When RL is clearly the right answer

Cheap simulation (games, robot simulators), a reward that comes for free (score, win rate, profit), a clearly defined action space, and a solution space too large to cover with hand-written rules — meet two or more of these, and RL is usually worth a shot. See the project suggestions in Learning Paths.

The Environment: Choosing, Building, Wrapping ​

The environment is the single biggest source of risk in an RL project. It determines how many samples a day of training buys you, whether your evaluation results are trustworthy, and whether the learned policy transfers at all.

1. Three Sources ​

SourceBest forCostRisk
Off-the-shelf environment libraryThe project is fundamentally about "running algorithms" (learning, reproduction, comparison)LowFew environment bugs, but the env may not match your business
Off-the-shelf + customizationYour business closely resembles a public environment (e.g., scheduling ≈ job-shop)MediumThe custom parts introduce new bugs
Fully customUnique to your business (recommendation, trading, physical systems)HighThe environment is the product — every bug hides here

For choosing among environment libraries, see the Datasets & Tools Index: Gymnasium is the de facto standard API, MuJoCo/Brax cover continuous control (and Brax runs on GPU), while Isaac-style platforms handle robotics. Whichever you pick, the very first task is a 30-line "smoke test" for the environment: run 1,000 steps with random actions and verify that step's return shapes, the conditions that trigger done, and the reward range all match expectations.

2. Three Disciplines for Building Your Own Environment ​

A custom environment (a recommendation simulator, a scheduling simulation, and so on) is exactly what "the environment is the product" refers to. Three disciplines are non-negotiable:

  1. Define the interface before writing the logic. Subclass gymnasium.Env directly and implement reset, step, observation_space, and action_space. For Gymnasium API details, see the Progressive Tutorial.
  2. Declare observations and actions explicitly with spaces and validate with space.contains(). Getting the declaration wrong (e.g., a discrete action written as continuous) is a top cause of training failing silently.
  3. Keep the environment separate from the training code. The environment is a standalone process or service, and the training code only talks to it through the interface. That way you can swap implementations (Python → C++ → distributed) without ever touching the algorithm code.

3. A Minimal Custom Environment ​

Below is a toy "1D cart moves toward a goal" environment that shows the interface skeleton:

python
import gymnasium as gym
import numpy as np
from gymnasium import spaces

class ReachGoalEnv(gym.Env):
    """1D cart motion: the state is position and velocity, the action is a
    force, and the goal is to come to rest near [0, 0]."""

    metadata = {"render_modes": []}

    def __init__(self):
        super().__init__()
        # Action: a continuous force in [-1, 1]
        self.action_space = spaces.Box(low=-1.0, high=1.0, shape=(1,), dtype=np.float32)
        # Observation: position and velocity
        self.observation_space = spaces.Box(low=-np.inf, high=np.inf, shape=(2,), dtype=np.float32)
        self._state = np.zeros(2, dtype=np.float32)
        self._step_count = 0

    def reset(self, *, seed=None, options=None):
        super().reset(seed=seed)          # required: seeds np_random
        self._state = np.array([self.np_random.uniform(-1, 1), 0.0], dtype=np.float32)
        self._step_count = 0
        return self._state.copy(), {}

    def step(self, action):
        action = np.clip(action, -1.0, 1.0).item()
        pos, vel = self._state
        vel = vel + 0.1 * action           # simple dynamics
        pos = pos + vel
        self._state = np.array([pos, vel], dtype=np.float32)
        self._step_count += 1

        # Reward: closer to the goal is better; also encourages coming to rest
        reward = -abs(pos) - 0.5 * abs(vel)
        # Termination: near the goal or out of steps
        terminated = abs(pos) < 0.05 and abs(vel) < 0.05
        truncated = self._step_count >= 200
        info = {"dist": abs(pos)}
        return self._state.copy(), reward, terminated, truncated, info

Memorize the interface conventions

In current Gymnasium, reset returns (obs, info) and step returns (obs, reward, terminated, truncated, info), and terminated (task completed) must be kept separate from truncated (cut short by external forces, such as a time limit). Treating a timeout as "task failed" or "task succeeded" systematically distorts the learning signal. See the tutorial's pitfall list.

4. Environment Acceptance Checklist ​

Before a custom environment is signed off, walk through every item:

  • [ ] 1,000 rollouts under a random policy raise no exceptions; observation/reward types match the space declarations
  • [ ] Rewards are bounded (or at least have bounded variance), so reward normalization can be designed later
  • [ ] reset produces the same initial-state sequence for the same seed
  • [ ] A simple hand-written heuristic scores points, confirming the environment is solvable (i.e., the environment itself isn't impossible)
  • [ ] step throughput has been measured (steps per second) to confirm it supports the target sample budget

Steps 2 and 3: Reward Design and Baselines First ​

1. Design the Reward Before Picking the Algorithm ​

"The reward is the spec" — the reward function is the only machine-readable expression of the product requirements. Get this step wrong and everything downstream is wasted; see Reward Engineering for the full treatment. At the practice level there's exactly one rule: write a reward design document first — for every reward term, list its source, its value, why it's set that way, and how it might be gamed. Only then start writing code.

2. Baselines First: No Baselines, No Algorithm ​

Many projects fire up SAC or PPO on day one, the agent doesn't learn, and nobody can tell whether the algorithm or the environment is at fault. The right move is to set up a ladder of difficulty:

LevelPolicyPurposeExpected outcome
0Random policyVerify the environment responds and rewards flowEstablishes the "floor"
1Hand-written heuristic (rules / greedy / linear)Verify the problem is solvable at allEstablishes a "reference line"
2Supervised warm-up (behavior cloning / supervised baseline)Reach a reasonable level quickly when offline data existsA "cheap upper bound"
3Simple RL (DQN / REINFORCE)Verify the RL signal path worksAt least matches the heuristic
4Full RL (PPO / SAC + tuning)Push toward the ceilingSignificantly beats the heuristic

Why run the heuristic first

A heuristic is a ruler: it tells you directly whether the environment's reward signal is dense enough to learn from. If the heuristic scores 80 with ease while RL sits at 30, the problem almost certainly lies in the reward or observation design — not the algorithm.

An all-too-common mistake is skipping Level 1, going straight to PPO, and spending two weeks tuning hyperparameters before discovering the hand-written rule was better all along. Time saved at Level 1 always gets repaid with interest in Steps 4-6.

The Minimal Training Loop ​

With the environment and baselines in place, it's time to build the training loop. The principle is minimal skeleton, then incremental polish: the first version only needs to do three things — run, save the model, and produce a learning curve. Every optimization (vectorization, parallelism, distributed training) waits until the signal path is confirmed correct.

1. The Minimal Skeleton: A Framework-Agnostic Loop ​

python
"""Minimal RL training-loop skeleton: depends only on gymnasium + numpy.
Goal: run end to end, save the model, produce a curve. Swapping algorithms
does not change the skeleton's structure."""

import gymnasium as gym
import numpy as np

class RandomPolicy:
    """Placeholder policy: swap in any algorithm exposing .act(obs)."""
    def __init__(self, env):
        self.env = env
    def act(self, obs):
        return self.env.action_space.sample()

def collect_episode(env, policy, max_steps=500):
    """Run one episode, return (total return, number of steps)."""
    obs, _ = env.reset()
    total_reward, steps = 0.0, 0
    for _ in range(max_steps):
        action = policy.act(obs)
        obs, reward, terminated, truncated, _ = env.step(action)
        total_reward += reward
        steps += 1
        if terminated or truncated:
            break
    return total_reward, steps

def train(env_id, num_episodes=2000, seed=0):
    env = gym.make(env_id)
    # 1. Seed the environment (seed every environment)
    env.reset(seed=seed)
    policy = RandomPolicy(env)

    returns = []
    for ep in range(num_episodes):
        # 2. Re-seed the environment every episode for reproducibility
        env.reset(seed=seed + ep)
        total_reward, _ = collect_episode(env, policy)
        returns.append(total_reward)

        # 3. Logging: print the window mean every 100 episodes (coarse first)
        if (ep + 1) % 100 == 0:
            mean = float(np.mean(returns[-100:]))
            print(f"episode {ep+1:>5d} | mean_return(100) = {mean:.2f}")

    # 4. Save the model + the curve
    np.save("returns.npy", np.array(returns))
    print("done, returns saved to returns.npy")
    env.close()

if __name__ == "__main__":
    train("CartPole-v1")

Three design decisions in this skeleton are worth remembering:

  1. The policy object exposes a single interface, act(obs) — DQN, REINFORCE, and PPO can all plug in the same way, so swapping algorithms barely touches the training loop. For the algorithm internals, see Value-Based Learning and Policy Gradient Methods.
  2. Every episode re-seeds reset — this separates environment randomness from algorithm randomness, so when something goes wrong you know which layer to blame.
  3. Logging starts coarse and gets finer — the first version prints only a windowed mean; once the signal checks out, add diagnostics like mean TD error, entropy, and KL gradually (see Tuning in Practice).

2. From Skeleton to Production: Three Stages ​

StageWhat you doWhen
SkeletonSingle process, step-by-step episodes; confirm learning happensDay one
VectorizedParallel collection with gymnasium.vector or SubprocVecEnvWhen single process is too slow
FrameworkRewrite with, or port to, SB3 / CleanRL / RLlibOnce the skeleton proves it works

Don't rush to a framework

Write a 50-line training loop by hand first, then go read the framework comparison. The most common outcome of reaching for a framework immediately is that the algorithm runs, but you don't understand a single number in the logs, and when a bug hits you have no idea where to start.

Evaluation and Logging ​

Decide the evaluation protocol at the start of the project, not after training — because every "this looks better" decision you make during training rests on an evaluation judgment. The full protocol lives in Evaluation in Practice; here are the three minimal pieces that actually work:

  1. A fixed evaluation function: every N environment steps during training, run num_eval_episodes episodes with fixed seeds and log "mean eval return ± interquartile range," kept separate from training returns.
  2. Persist the learning curves: append (timestep, train_return, eval_return, seed) to a CSV on every run — this is the raw material for all later diagnosis.
  3. Write an experiment record for every run: config (hyperparameters, seed, environment), git commit, log path. An unrecorded experiment is an experiment that didn't happen.

The core concepts behind evaluation — sample efficiency vs. final performance, multi-seed matrices, IQR — come from Evaluation and Benchmarks; the practice-level template is in Building an RL Evaluation Suite from Scratch.

Hardening Reproducibility ​

Reproducibility in RL is an order of magnitude worse than in supervised learning: identical code and hyperparameters can produce different curves on a different machine, a different GPU, or even a different NumPy version. Here are the countermeasures, ranked by ROI:

1. Three-Layer Seed Management ​

LayerWhat you seedCode
Globalrandom, numpy, torchrandom.seed(seed); np.random.seed(seed); torch.manual_seed(seed)
EnvironmentGymnasium environmentsenv.reset(seed=s); env.action_space.seed(s)
Inside the algorithmReplay sampling, noise, action samplingThe framework's built-in seed(s) (e.g., SB3's PPO(..., seed=...))

A global seed is not a silver bullet

Some PyTorch operators are not deterministic on GPU (anything built on atomicAdd-style operations), and multi-GPU or parallel sampling makes bit-level reproduction flatly impossible. The pragmatic goal isn't "reproduce bit for bit" but reproduce the performance distribution for a given config (same mean and IQR) — and that requires multiple seeds, not one. See the seed pitfalls in Tuning in Practice.

2. Configuration Management: Hydra ​

Hyperparameters scattered through the code are the biggest hygiene problem in RL projects. Use a configuration system to separate "experiments" from "code":

yaml
# config/train.yaml (Hydra config example)
seed: 42
env_id: CartPole-v1

algorithm:
  name: ppo
  learning_rate: 3.0e-4
  gamma: 0.99
  gae_lambda: 0.95
  clip_range: 0.2
  ent_coef: 0.0
  n_steps: 2048
  batch_size: 64

evaluation:
  eval_episodes: 10
  eval_freq: 10000   # evaluate every N environment steps
  seeds: [0, 1, 2]   # multiple evaluation seeds

Override any field from the command line (python train.py algorithm.learning_rate=1e-3), and archive the full config of every experiment alongside its logs. "Reproducing a result" then degenerates into "re-running a config."

3. Locking Dependencies and Versions ​

  • Pin requirements.txt / pyproject.toml to minor versions (RL code couples tightly to specific NumPy/Gymnasium versions).
  • Record the versions of gymnasium, numpy, torch, and stable-baselines3 in every experiment record.
  • If you can, freeze a Docker image with docker build, or use the lockfiles of uv/poetry.

Deployment and Monitoring ​

Training a good policy is only half the project. The classic failure mode at launch is: great offline scores, terrible online performance. The usual root causes are distribution shift and a mismatch between evaluation and deployment.

1. Serving Patterns for Policies ​

ScenarioServing patternExamples
Offline batch decisionsBatch job run on a scheduleProduction scheduling, batch recommendation
Online single decisionsLow-latency inference serviceAd bidding, real-time control
Online environment interactionThe environment lives in the business system; the policy is called backRobotics, dialogue systems

The universal practice: export the policy to an inference-only format (ONNX / TorchScript / JAX), so the serving stack depends only on the "observation → action" forward function and carries no training dependencies whatsoever. The training loop, the replay buffer, and the optimizer have no business being in the serving path.

2. The Monitoring Trinity ​

  1. Action distribution: drift in the mean/variance/spread of online actions means the observation distribution has shifted.
  2. Observation distribution: compute statistics on live observations (means, quantiles, approximate hashes) and compare them against the training distribution.
  3. Business metrics: rewards often can't be measured online directly (delayed rewards, human judgment), so proxy business metrics must be watched at the same time.
text
Drift detection flow (weekly patrol):
 collect this week's live observations → compare distributions against
 the training data (KS test / histograms)
 → significant difference? → yes: trigger a retrain evaluation;
   no: keep watching

3. Rollback and Safety Rails ​

  • Version every policy, and keep the last N versions in the model server so you can roll back with one click.
  • Before going live, run a shadow deployment: the policy only logs its actions without executing them, and you compare against the existing system for a week.
  • Keep a rule-based fallback on the business side: when a policy action is abnormal (out of range, NaN, timeout), fall back to the rule policy. Safety-related discussion is in RL Design Principles and Common Pitfalls.

Iteration Rhythm: Make It Run First, Then Tune ​

To close, here's an executable rhythm sheet. Most failed RL projects share one trait — not "the algorithm didn't work," but tuning before the pipeline actually ran, or tuning without discipline after it did.

PhaseSuggested durationDeliverableExit condition
Steps 0-21-2 weeksDecision checklist, environment acceptance sign-off, reward design docEnvironment runs; heuristic scores
Steps 3-43-5 daysSkeleton training loop, random/heuristic baseline curvesCan save models and produce learning curves
Steps 5-63-5 daysEvaluation script, CSV logs, Hydra configsOne command fully reproduces an experiment
Step 7In parallel with trainingDeployment pipeline, monitoring and alertsShadow deployment works
Tuning loopWeeklyChange one variable per round; keep experiment recordsBusiness acceptance criteria met

The one-loop-a-day discipline

Treat "change code → run experiment → read curves → write down conclusions" as a single loop and complete at least one loop per day. RL experiments routinely take hours, and the nightmare scenario is changing five variables, running for five hours, and ending up unable to say which variable did anything. Single-variable experiments, make-it-run-before-tuning, and experiment records — these three habits alone will save you 60% of your total project time.

Further Reading ​

References ​

  • Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. MIT Press. (The whole book, especially Chapters 1, 2, and 8 on the agent-environment interface and tabular methods; free online edition at incompleteideas.net/book/the-book-2nd.html)
  • Gymnasium official documentation (Farama Foundation): gymnasium.farama.org; Env interface reference at gymnasium.farama.org/api/env/
  • Stable-Baselines3 official documentation: stable-baselines3.readthedocs.io
  • Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560 (the classic treatment of RL reproducibility)
  • Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133 (empirical evidence on seed sensitivity)
  • Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292 (why implementation details matter more than the paper's equations)
  • Irpan, A. (2018). Deep Reinforcement Learning Doesn't Work Yet. alexirpan.com/2018/02/14/rl-hard.html (the famous blog post on why RL projects fail)
  • Hydra official documentation: hydra.cc