Skip to content

Anatomy of an RL System

On this page The six layers of a complete RL system — environment, experience collection, learning algorithm, policy, evaluator, and safety guardrails — and why "the environment is the product": 80% of RL project cost sits in the environment.

Anatomy of an RL System ​

In one sentence: this page takes "an RL project" apart into a machine you can inspect and repair — six layers: environment, experience collection, learning algorithm, policy, evaluator, and safety guardrails. After reading it, you'll look at any RL codebase (Stable-Baselines3, RLlib, someone else's GitHub project) and ask not "what file is this?" but "which layer does this blob belong to, and what's its input-output contract?" More importantly, you'll come to understand the industry saying "the environment is the product" — why the biggest share of an RL project's cost sits not in the algorithm but in the environment.

1. Overview: The Six-Layer Architecture ​

text
                       ┌──────────────────────────────┐  ← limit · veto · rollback
                       │  ⑥ Guardrails                │
                       │  limit · veto · rollback     │
                       └──────────────────────────────┘
   ┌────────────────────────────────────────────────────────────┐
   │  ① Environment (real world / simulator / API / users)      │
   │     state s ──▶ ──▶ action a · reward r ──▶               │
   └────────────────────────────────────────────────────────────┘
              ▲                                                │
              │  (s, a, r, s') tuples                          │ actions
   ┌──────────┴───────────────┐        ┌──────────────────────────▼┐
   │  ② Experience            │        │  ④ Policy                 │
   │     Collection           │        │     π(a|s) snapshots /    │
   │     replay buffer /      │        │     versions              │
   │     trajectory pool      │        │                           │
   └──────────┬───────────────┘        └───────────────────────────┬┘
              │ training batches                                   │ parameter sync
   ┌──────────▼───────────────┐        ┌───────────────────────────┴┐
   │  ③ Learner               │        │  ⑤ Evaluator              │
   │     gradient updates     │        │     offline metrics /     │
   │     (PPO / SAC / …)      │        │     live monitoring       │
   └──────────────────────────┘        └───────────────────────────┘

Each layer solves one well-defined problem, and the layers connect to one another through interface contracts rather than shared global state:

LayerInputOutputThe key question of this layer
① Environmentaction astate s′, reward r, done flagIs the simulation realistic enough and fast enough? Is the reward wired correctly?
② Experience collectionthe policy's stream of actionstrajectories / (s,a,r,s′) batchesParallelize sampling? Prioritized buffer?
③ Learning algorithmtraining batchesupdated parametersAre the gradients stable? Is sample efficiency high?
④ Policystate s (real-time)action a (low latency)Serving, version management, inference latency
⑤ Evaluatorenvironment replays / logslearning curves, task success rate, live metricsIs the evaluation protocol fair? Are the metrics real?
⑥ Guardrailsbehavior of all layersinterception / rollback / degradation signalsHow to block out-of-bounds actions, and what to do when things go wrong

What the six-layer architecture is for

This page isn't something to memorize — it's an engineering checklist. While building, ask yourself layer by layer: "Is this layer's interface defined? Tested? If something breaks, can I pinpoint it here?" Want to build one hands-on? The eight-step pipeline in Build an RL Project from Scratch is essentially the six-layer architecture, implemented.

2. Layer by Layer ​

1. The Environment Layer: the RL "World" ​

The environment defines the problem itself: state space, action space, transition rules, reward signal, termination conditions. It may be a simulator (MuJoCo, Isaac, Gymnasium environments), a real physical system, or an API that "treats the user as the environment."

The three most common forms of the environment layer:

text
Real environment     Simulated environment    API environment
Physical world       Numerical simulator      Recommendation / ads / web
Expensive samples    Cheap samples            Delayed feedback
No replays           Replayable               A/B-testable

The environment's interface contract is the one Gymnasium set: env.step(action) → (obs, reward, done, info). Nearly every framework honors it, and it's the subject of the first lesson in the Progressive Gymnasium Tutorial. For environment selection and environment-building guides, see Datasets & Tools.

The hidden cost of environments

An environment is not "call a library and be done with it." A production-grade environment must answer: What are the observation dimensions and types? Where does the reward signal come from, and how noisy is it? How are termination conditions determined? Are random seeds controllable? Can it replay? The answers directly determine whether the algorithm can learn and whether evaluation can be reproduced. This layer is the load-bearing wall for everything built on top.

2. The Experience Collection Layer ​

RL's training data doesn't come ready-made — it's collected. This layer is responsible for:

  • Collection: the policy runs in the environment, producing (s, a, r, s′) tuples or full trajectories;
  • Storage: an experience replay buffer (required by the DQN family; see Value-Based Learning) or a short-term on-policy trajectory pool;
  • Parallelism: multi-environment parallel sampling (vectorized environments) is standard equipment in deep RL — A2C/PPO's n_envs and the worker processes of distributed RL all live in this layer;
  • Sampling: in what order experiences are fed to the learner (uniform random / prioritized replay / latest-first).
text
Sampling throughput = environment speed × parallelism
  · CPU environments: multi-process parallelism (e.g., 16 CartPole instances)
  · GPU environments: Brax packs tens of thousands of environments onto one card
  · Real environments: the bottleneck is usually the environment itself — you take what you can get

Why optimize this layer first

80% of RL debugging time goes to "training isn't moving," and "training isn't moving" is usually slow sampling or poor data quality (e.g., the environment keeps returning the same state). Get sampling throughput and buffer logic measured before touching the algorithm. For the engineering reality of parallel sampling and GPU acceleration, see Choosing Frameworks & Tools.

3. The Learning Algorithm Layer (Learner) ​

This is what most people mean by "the RL itself": given a training batch, update the parameters of the policy or value function. Within the layer, the subdivisions are:

Algorithm familyWhat it learnsRepresentativesUpdate signal
Value-basedQ(s,a) / V(s)DQN, RainbowTD error (replay + target network)
Policy gradientπ(a|s) directlyREINFORCE, PPOadvantage A × log-probability gradient
Actor-criticpolicy + value, two networksA2C, PPO, SAC, TD3the value net supplies the advantage; the policy net updates with it
Model-basedworld model + planningDreamer, TD-MPCprediction error + planning returns

The selection logic — which family for which scenario — has a full genealogy table in The Actor-Critic Family. The learning layer also has a hidden role: hyperparameters. PPO's clip, SAC's temperature coefficient, and learning-rate schedules all take effect here — and RL is an order of magnitude more sensitive to them than supervised learning. Remedies are in Hyperparameter Tuning in Practice.

4. The Policy Layer: the Face of Online Serving ​

Once trained — or even while training — the policy needs to serve online. The engineering problems of this layer are often overlooked by beginners:

  • Version management: save a checkpoint every N steps; which version runs online, and how do you roll back?
  • Inference latency: real-time decisions (trading, recommendation) are latency-sensitive — how large can the network be and still fit the latency budget?
  • Serving: how does the policy model talk to the online system (e.g., wrapping the Q-network as an RPC service)?
  • Drift monitoring: what if the online state distribution no longer matches the training distribution?

The policy layer is often underestimated because "it's just one forward pass" — yet it's exactly this layer that decides whether your system merely "runs" or is actually "deployable."

Online policy ≠ training policy

In training, the policy carries exploration noise (ε-greedy, entropy regularization); online it usually executes greedily with the noise removed. This behavior-policy vs. target-policy distinction — and the "offline evaluation is hard" problem it creates — is the first hurdle of putting RL into production; see Offline Reinforcement Learning.

5. The Evaluator Layer: RL's "Quality Assurance Department" ​

RL has no ready-made test set, and the evaluator layer exists to answer "is this policy actually any good?":

  • In-training evaluation: periodically replay environments with fixed seeds, recording mean/median returns, success rates, and learning curves;
  • Comparative evaluation: multiple algorithms, multiple seeds, fixed budget — producing a fair comparison table;
  • Online evaluation: live metrics (click-through rate, task success rate), A/B tests, drift monitoring.

The technical details of evaluation protocols (why a single seed isn't enough, sample efficiency vs. final performance, reporting standards) are covered on Evaluation & Benchmarks and Building an RL Evaluation from Scratch. Here we stress just one iron rule:

A project without an evaluation protocol is a delivery without acceptance criteria. Learning curves, multi-seed variance, a fixed evaluation environment and budget — none of these are optional. Without them, you can't answer the question "so what do your experiments actually show?" — whether it comes from an interviewer or from yourself.

6. The Guardrails Layer: the Last Line of Defense ​

RL agents "try things on their own," so production systems must add guardrails at the behavior level:

text
Three typical guardrails
  Action limiting    action leaves the safe range → clamp / veto / fall back to a conservative policy
  Constraint checks  constraint violated (e.g., collision, limit exceeded) → terminate the episode and penalize
  Human fallback     critical scenarios (oversized trade, robot approaching a human) → force a switch to human control

The philosophy of guardrails: RL is in charge of being smart; guardrails are in charge of not causing trouble. Safety constraints in autonomous driving, business guardrails in recommendation, refusals in large language models — all are concrete forms of this layer; see Autonomous Driving Decision-Making. This also echoes the selection criterion in RL vs. Neighboring Paradigms — "can you afford the cost of trial and error?": guardrails are the engineering backstop for that cost.

Guardrails can't fix the reward

What guardrails intercept is "behavior out of bounds" — they cannot stop "rewards being gamed" (reward hacking). A robot may farm cleaning rewards entirely within the safety boundary. The design of the reward itself must be solved by Reward Engineering; a guardrail is only the last physical gate. The two are complementary, not substitutes.

3. Three Data Flows: Online vs. Offline vs. Simulated ​

The six-layer architecture doesn't change, but "where the experience comes from" determines the entire engineering shape of the system. Comparing the three data flows:

DimensionOnline RL (live sampling)Offline RL (historical data)Simulated RL (simulator)
Data sourcecurrent policy interacts with the environment in real timeexisting logs; no interaction during traininggenerated on demand by a simulator
Exploration freedomhighnone (only actions present in the logs)extremely high (replayable, parallelizable)
Sample costhigh (real systems are expensive)medium (collected once)low (compute traded for samples)
Distribution drift riskyes (training changes behavior)low (data fixed) but severe OOD problemyes (simulation ≠ reality; the Sim2Real gap)
Main challengessampling speed, exploration, instabilityOOD actions and value overestimationsimulation fidelity, transfer
Representative scenariosgame training, sim trainingrecommendation logs, financial history, robot replaysrobotics, autonomous driving, games

The engineering implications of each:

  1. For online RL, the bottleneck is "sampling and training must form a fast closed loop" — hence asynchronous architectures (A3C's parallel sampling), vectorized environments, and distributed samplers;
  2. For offline RL, the bottleneck is on the data side — OOD actions give the value network "bootstrapping hallucinations" (overestimation); solutions and their limits are in Offline Reinforcement Learning;
  3. For simulated RL, the bottleneck is "the gap between simulation and reality" — domain randomization and Sim2Real are the bridging measures; see Robot Control and Sim2Real.
text
The most common best-practice pipeline:
  simulated RL trains the policy → offline data accumulates → offline RL fine-tuning → go live (small online updates)
  Every step lowers the real-world cost of trial and error — the mainstream path of RL engineering.

4. "The Environment Is the Product": Why 80% of the Cost Lives in the Environment ​

This is the most widely quoted saying in RL engineering circles. Unpacked, it carries four layers of meaning:

1. The Environment Decides "Whether the Problem Can Be Solved at All" ​

Algorithms are general; the environment is the problem incarnate. The same PPO converges in two hours in a correctly written environment and never learns in a buggy one (wrong rewards, leaked state, misjudged done flags). A silent bug in the environment costs an order of magnitude more than choosing the wrong algorithm — this is exactly the origin of the "environment bugs silently sabotage learning" entry in Common Pitfalls & Anti-Patterns.

2. The Environment Holds the Lifeline of Data Efficiency ​

RL is notorious for poor sample efficiency: supervised learning gets by with 10,000 samples, while RL routinely needs a million steps. And samples come from the environment — so the environment's speed (FPS) multiplies straight into training cost:

text
Training cost ≈ environment speed⁻¹ × steps required × unit compute price
  · Making the environment 10× faster (swap simulators, parallelize) = 10× lower training cost
  · That is why industry spends so much effort "making the environment fast and correct"

3. The Environment Decides "Whether Evaluation Can Be Reproduced" ​

The evaluator replays on the environment. An irreproducible environment (sloppy random-seed management, hidden state) → irreproducible evaluation → every experimental conclusion in doubt. This is why Evaluation & Benchmarks lists "fixed seeds, multiple runs" as the first element of an evaluation protocol.

4. The Environment Decides "Whether the Reward Is Written Correctly" ​

The reward signal belongs to the environment layer. Reward hacking, reward sparsity, objective drift — every problem discussed under Reward Engineering must ultimately be solved in the environment layer's reward function. The ultimate meaning of "the environment is the product": what you deliver to the business is not an algorithm but the precise specification of the problem — "environment + reward + evaluation" — and the algorithm is just a solver for that specification.

A counterintuitive corollary

If "the environment is the product," then the first deliverable of an RL project should be environment documentation plus environment tests (state space, action space, reward formula, termination conditions, reproducibility) — not a model. Write the environment spec before the algorithm; this order steers you clear of most "it won't learn" pits.

5. Comparing with the Anatomy of an ML System ​

The typical "anatomy of an ML system" has a four-piece kit: data pipeline, model, loss function, and evaluation set. Mapping RL's six layers onto it:

ML systemRL systemAnalogy and difference
Data pipeline (feature engineering)environment + experience collectionboth "feed data," but RL's data is generated by the policy itself and must form a closed loop
Model (forward inference)policy + learning algorithmboth are parameterized mappings, but RL training depends on its own sampling
Loss function (supervisory signal)reward function + TD/advantage objectivesa supervised loss is fixed explicitly; RL's "loss" is dynamic, determined by trajectories
Evaluation set (fixed test set)evaluator (replay environments / protocols)an ML test set is frozen; RL evaluation faces environment randomness and multi-seed variance
(usually no guardrails)safety guardrailsRL's behavior-generating nature makes guardrails mandatory

This comparison shows that an RL system is not a "harder ML system" but a closed-loop system: every stage depends on the runtime output of the others. It also explains why RL troubleshooting chains run so long — a bug can hide in the environment, sampling, buffer, gradients, or evaluation, and the layers amplify one another. The layer-by-layer troubleshooting checklist is in Common Pitfalls & Anti-Patterns.

6. Into the Practice Module ​

Now that you've seen the six-layer architecture, let's make it spin — this site's practice module is the layer-by-layer landing of it:

One last suggestion

Save the six-layer diagram from this page as the "architecture sketch" starting point for every RL project. Even if your project is a few dozen lines of script, it's worth marking in the comments "this is the environment layer, this is sampling, this is learning, this is evaluation" — the habit turns you from "someone who just calls libraries" into "someone who builds systems." That's a watershed on your resume and in interviews.

Further Reading ​

References ​

  • Farama Foundation. Gymnasium Documentation. https://gymnasium.farama.org/ — the official standard for the environment layer's interface contract (reset/step/render, observation and action spaces).
  • Stable-Baselines3 Contributors. Stable-Baselines3 Documentation. https://stable-baselines3.readthedocs.io/ — production-ready, off-the-shelf implementations of the learning-algorithm layer.
  • Ray Project. Ray RLlib Documentation. https://docs.ray.io/en/latest/rllib/index.html — the representative implementation of parallelization and distributed training for the experience-collection layer.
  • Weng, L. (2018). The Challenges of Reinforcement Learning. https://lilianweng.github.io/posts/2018-02-19-rl-overview/ — a systematic discussion of "what makes RL engineering hard," echoing this page's "the environment is the product" argument.
  • Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. https://arxiv.org/abs/1709.06560 — required reading for the evaluator layer: the statistical traps of reproducibility and fair comparison in RL experiments.
  • Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.), §1.1–1.3. MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the authoritative source for the agent-environment interface concept.