Appearance
Anatomy of an RL System
In one sentence: this page takes "an RL project" apart into a machine you can inspect and repair — six layers: environment, experience collection, learning algorithm, policy, evaluator, and safety guardrails. After reading it, you'll look at any RL codebase (Stable-Baselines3, RLlib, someone else's GitHub project) and ask not "what file is this?" but "which layer does this blob belong to, and what's its input-output contract?" More importantly, you'll come to understand the industry saying "the environment is the product" — why the biggest share of an RL project's cost sits not in the algorithm but in the environment.
1. Overview: The Six-Layer Architecture
text
┌──────────────────────────────┐ ← limit · veto · rollback
│ ⑥ Guardrails │
│ limit · veto · rollback │
└──────────────────────────────┘
┌────────────────────────────────────────────────────────────┐
│ ① Environment (real world / simulator / API / users) │
│ state s ──▶ ──▶ action a · reward r ──▶ │
└────────────────────────────────────────────────────────────┘
▲ │
│ (s, a, r, s') tuples │ actions
┌──────────┴───────────────┐ ┌──────────────────────────▼┐
│ ② Experience │ │ ④ Policy │
│ Collection │ │ π(a|s) snapshots / │
│ replay buffer / │ │ versions │
│ trajectory pool │ │ │
└──────────┬───────────────┘ └───────────────────────────┬┘
│ training batches │ parameter sync
┌──────────▼───────────────┐ ┌───────────────────────────┴┐
│ ③ Learner │ │ ⑤ Evaluator │
│ gradient updates │ │ offline metrics / │
│ (PPO / SAC / …) │ │ live monitoring │
└──────────────────────────┘ └───────────────────────────┘Each layer solves one well-defined problem, and the layers connect to one another through interface contracts rather than shared global state:
| Layer | Input | Output | The key question of this layer |
|---|---|---|---|
| ① Environment | action a | state s′, reward r, done flag | Is the simulation realistic enough and fast enough? Is the reward wired correctly? |
| ② Experience collection | the policy's stream of actions | trajectories / (s,a,r,s′) batches | Parallelize sampling? Prioritized buffer? |
| ③ Learning algorithm | training batches | updated parameters | Are the gradients stable? Is sample efficiency high? |
| ④ Policy | state s (real-time) | action a (low latency) | Serving, version management, inference latency |
| ⑤ Evaluator | environment replays / logs | learning curves, task success rate, live metrics | Is the evaluation protocol fair? Are the metrics real? |
| ⑥ Guardrails | behavior of all layers | interception / rollback / degradation signals | How to block out-of-bounds actions, and what to do when things go wrong |
What the six-layer architecture is for
This page isn't something to memorize — it's an engineering checklist. While building, ask yourself layer by layer: "Is this layer's interface defined? Tested? If something breaks, can I pinpoint it here?" Want to build one hands-on? The eight-step pipeline in Build an RL Project from Scratch is essentially the six-layer architecture, implemented.
2. Layer by Layer
1. The Environment Layer: the RL "World"
The environment defines the problem itself: state space, action space, transition rules, reward signal, termination conditions. It may be a simulator (MuJoCo, Isaac, Gymnasium environments), a real physical system, or an API that "treats the user as the environment."
The three most common forms of the environment layer:
text
Real environment Simulated environment API environment
Physical world Numerical simulator Recommendation / ads / web
Expensive samples Cheap samples Delayed feedback
No replays Replayable A/B-testableThe environment's interface contract is the one Gymnasium set: env.step(action) → (obs, reward, done, info). Nearly every framework honors it, and it's the subject of the first lesson in the Progressive Gymnasium Tutorial. For environment selection and environment-building guides, see Datasets & Tools.
The hidden cost of environments
An environment is not "call a library and be done with it." A production-grade environment must answer: What are the observation dimensions and types? Where does the reward signal come from, and how noisy is it? How are termination conditions determined? Are random seeds controllable? Can it replay? The answers directly determine whether the algorithm can learn and whether evaluation can be reproduced. This layer is the load-bearing wall for everything built on top.
2. The Experience Collection Layer
RL's training data doesn't come ready-made — it's collected. This layer is responsible for:
- Collection: the policy runs in the environment, producing
(s, a, r, s′)tuples or full trajectories; - Storage: an experience replay buffer (required by the DQN family; see Value-Based Learning) or a short-term on-policy trajectory pool;
- Parallelism: multi-environment parallel sampling (vectorized environments) is standard equipment in deep RL — A2C/PPO's
n_envsand the worker processes of distributed RL all live in this layer; - Sampling: in what order experiences are fed to the learner (uniform random / prioritized replay / latest-first).
text
Sampling throughput = environment speed × parallelism
· CPU environments: multi-process parallelism (e.g., 16 CartPole instances)
· GPU environments: Brax packs tens of thousands of environments onto one card
· Real environments: the bottleneck is usually the environment itself — you take what you can getWhy optimize this layer first
80% of RL debugging time goes to "training isn't moving," and "training isn't moving" is usually slow sampling or poor data quality (e.g., the environment keeps returning the same state). Get sampling throughput and buffer logic measured before touching the algorithm. For the engineering reality of parallel sampling and GPU acceleration, see Choosing Frameworks & Tools.
3. The Learning Algorithm Layer (Learner)
This is what most people mean by "the RL itself": given a training batch, update the parameters of the policy or value function. Within the layer, the subdivisions are:
| Algorithm family | What it learns | Representatives | Update signal |
|---|---|---|---|
| Value-based | Q(s,a) / V(s) | DQN, Rainbow | TD error (replay + target network) |
| Policy gradient | π(a|s) directly | REINFORCE, PPO | advantage A × log-probability gradient |
| Actor-critic | policy + value, two networks | A2C, PPO, SAC, TD3 | the value net supplies the advantage; the policy net updates with it |
| Model-based | world model + planning | Dreamer, TD-MPC | prediction error + planning returns |
The selection logic — which family for which scenario — has a full genealogy table in The Actor-Critic Family. The learning layer also has a hidden role: hyperparameters. PPO's clip, SAC's temperature coefficient, and learning-rate schedules all take effect here — and RL is an order of magnitude more sensitive to them than supervised learning. Remedies are in Hyperparameter Tuning in Practice.
4. The Policy Layer: the Face of Online Serving
Once trained — or even while training — the policy needs to serve online. The engineering problems of this layer are often overlooked by beginners:
- Version management: save a checkpoint every N steps; which version runs online, and how do you roll back?
- Inference latency: real-time decisions (trading, recommendation) are latency-sensitive — how large can the network be and still fit the latency budget?
- Serving: how does the policy model talk to the online system (e.g., wrapping the Q-network as an RPC service)?
- Drift monitoring: what if the online state distribution no longer matches the training distribution?
The policy layer is often underestimated because "it's just one forward pass" — yet it's exactly this layer that decides whether your system merely "runs" or is actually "deployable."
Online policy ≠ training policy
In training, the policy carries exploration noise (ε-greedy, entropy regularization); online it usually executes greedily with the noise removed. This behavior-policy vs. target-policy distinction — and the "offline evaluation is hard" problem it creates — is the first hurdle of putting RL into production; see Offline Reinforcement Learning.
5. The Evaluator Layer: RL's "Quality Assurance Department"
RL has no ready-made test set, and the evaluator layer exists to answer "is this policy actually any good?":
- In-training evaluation: periodically replay environments with fixed seeds, recording mean/median returns, success rates, and learning curves;
- Comparative evaluation: multiple algorithms, multiple seeds, fixed budget — producing a fair comparison table;
- Online evaluation: live metrics (click-through rate, task success rate), A/B tests, drift monitoring.
The technical details of evaluation protocols (why a single seed isn't enough, sample efficiency vs. final performance, reporting standards) are covered on Evaluation & Benchmarks and Building an RL Evaluation from Scratch. Here we stress just one iron rule:
A project without an evaluation protocol is a delivery without acceptance criteria. Learning curves, multi-seed variance, a fixed evaluation environment and budget — none of these are optional. Without them, you can't answer the question "so what do your experiments actually show?" — whether it comes from an interviewer or from yourself.
6. The Guardrails Layer: the Last Line of Defense
RL agents "try things on their own," so production systems must add guardrails at the behavior level:
text
Three typical guardrails
Action limiting action leaves the safe range → clamp / veto / fall back to a conservative policy
Constraint checks constraint violated (e.g., collision, limit exceeded) → terminate the episode and penalize
Human fallback critical scenarios (oversized trade, robot approaching a human) → force a switch to human controlThe philosophy of guardrails: RL is in charge of being smart; guardrails are in charge of not causing trouble. Safety constraints in autonomous driving, business guardrails in recommendation, refusals in large language models — all are concrete forms of this layer; see Autonomous Driving Decision-Making. This also echoes the selection criterion in RL vs. Neighboring Paradigms — "can you afford the cost of trial and error?": guardrails are the engineering backstop for that cost.
Guardrails can't fix the reward
What guardrails intercept is "behavior out of bounds" — they cannot stop "rewards being gamed" (reward hacking). A robot may farm cleaning rewards entirely within the safety boundary. The design of the reward itself must be solved by Reward Engineering; a guardrail is only the last physical gate. The two are complementary, not substitutes.
3. Three Data Flows: Online vs. Offline vs. Simulated
The six-layer architecture doesn't change, but "where the experience comes from" determines the entire engineering shape of the system. Comparing the three data flows:
| Dimension | Online RL (live sampling) | Offline RL (historical data) | Simulated RL (simulator) |
|---|---|---|---|
| Data source | current policy interacts with the environment in real time | existing logs; no interaction during training | generated on demand by a simulator |
| Exploration freedom | high | none (only actions present in the logs) | extremely high (replayable, parallelizable) |
| Sample cost | high (real systems are expensive) | medium (collected once) | low (compute traded for samples) |
| Distribution drift risk | yes (training changes behavior) | low (data fixed) but severe OOD problem | yes (simulation ≠ reality; the Sim2Real gap) |
| Main challenges | sampling speed, exploration, instability | OOD actions and value overestimation | simulation fidelity, transfer |
| Representative scenarios | game training, sim training | recommendation logs, financial history, robot replays | robotics, autonomous driving, games |
The engineering implications of each:
- For online RL, the bottleneck is "sampling and training must form a fast closed loop" — hence asynchronous architectures (A3C's parallel sampling), vectorized environments, and distributed samplers;
- For offline RL, the bottleneck is on the data side — OOD actions give the value network "bootstrapping hallucinations" (overestimation); solutions and their limits are in Offline Reinforcement Learning;
- For simulated RL, the bottleneck is "the gap between simulation and reality" — domain randomization and Sim2Real are the bridging measures; see Robot Control and Sim2Real.
text
The most common best-practice pipeline:
simulated RL trains the policy → offline data accumulates → offline RL fine-tuning → go live (small online updates)
Every step lowers the real-world cost of trial and error — the mainstream path of RL engineering.4. "The Environment Is the Product": Why 80% of the Cost Lives in the Environment
This is the most widely quoted saying in RL engineering circles. Unpacked, it carries four layers of meaning:
1. The Environment Decides "Whether the Problem Can Be Solved at All"
Algorithms are general; the environment is the problem incarnate. The same PPO converges in two hours in a correctly written environment and never learns in a buggy one (wrong rewards, leaked state, misjudged done flags). A silent bug in the environment costs an order of magnitude more than choosing the wrong algorithm — this is exactly the origin of the "environment bugs silently sabotage learning" entry in Common Pitfalls & Anti-Patterns.
2. The Environment Holds the Lifeline of Data Efficiency
RL is notorious for poor sample efficiency: supervised learning gets by with 10,000 samples, while RL routinely needs a million steps. And samples come from the environment — so the environment's speed (FPS) multiplies straight into training cost:
text
Training cost ≈ environment speed⁻¹ × steps required × unit compute price
· Making the environment 10× faster (swap simulators, parallelize) = 10× lower training cost
· That is why industry spends so much effort "making the environment fast and correct"3. The Environment Decides "Whether Evaluation Can Be Reproduced"
The evaluator replays on the environment. An irreproducible environment (sloppy random-seed management, hidden state) → irreproducible evaluation → every experimental conclusion in doubt. This is why Evaluation & Benchmarks lists "fixed seeds, multiple runs" as the first element of an evaluation protocol.
4. The Environment Decides "Whether the Reward Is Written Correctly"
The reward signal belongs to the environment layer. Reward hacking, reward sparsity, objective drift — every problem discussed under Reward Engineering must ultimately be solved in the environment layer's reward function. The ultimate meaning of "the environment is the product": what you deliver to the business is not an algorithm but the precise specification of the problem — "environment + reward + evaluation" — and the algorithm is just a solver for that specification.
A counterintuitive corollary
If "the environment is the product," then the first deliverable of an RL project should be environment documentation plus environment tests (state space, action space, reward formula, termination conditions, reproducibility) — not a model. Write the environment spec before the algorithm; this order steers you clear of most "it won't learn" pits.
5. Comparing with the Anatomy of an ML System
The typical "anatomy of an ML system" has a four-piece kit: data pipeline, model, loss function, and evaluation set. Mapping RL's six layers onto it:
| ML system | RL system | Analogy and difference |
|---|---|---|
| Data pipeline (feature engineering) | environment + experience collection | both "feed data," but RL's data is generated by the policy itself and must form a closed loop |
| Model (forward inference) | policy + learning algorithm | both are parameterized mappings, but RL training depends on its own sampling |
| Loss function (supervisory signal) | reward function + TD/advantage objectives | a supervised loss is fixed explicitly; RL's "loss" is dynamic, determined by trajectories |
| Evaluation set (fixed test set) | evaluator (replay environments / protocols) | an ML test set is frozen; RL evaluation faces environment randomness and multi-seed variance |
| (usually no guardrails) | safety guardrails | RL's behavior-generating nature makes guardrails mandatory |
This comparison shows that an RL system is not a "harder ML system" but a closed-loop system: every stage depends on the runtime output of the others. It also explains why RL troubleshooting chains run so long — a bug can hide in the environment, sampling, buffer, gradients, or evaluation, and the layers amplify one another. The layer-by-layer troubleshooting checklist is in Common Pitfalls & Anti-Patterns.
6. Into the Practice Module
Now that you've seen the six-layer architecture, let's make it spin — this site's practice module is the layer-by-layer landing of it:
- Build a complete project from scratch: the eight-step pipeline maps onto the six layers — see Build an RL Project from Scratch;
- Hands-on with the environment layer: see the Progressive Gymnasium Tutorial and Datasets & Tools;
- Choosing at the learning-algorithm layer: see Choosing Frameworks & Tools (which engineering scale each of SB3 / RLlib / Tianshou / CleanRL / Brax fits);
- Standard operating procedure for the evaluator layer: see Building an RL Evaluation from Scratch;
- The cross-layer problems of guardrails and rewards: see Reward Engineering and Common Pitfalls & Anti-Patterns.
One last suggestion
Save the six-layer diagram from this page as the "architecture sketch" starting point for every RL project. Even if your project is a few dozen lines of script, it's worth marking in the comments "this is the environment layer, this is sampling, this is learning, this is evaluation" — the habit turns you from "someone who just calls libraries" into "someone who builds systems." That's a watershed on your resume and in interviews.
Further Reading
- Build an RL Project from Scratch — the eight-step pipeline that lands the six-layer architecture; the required next stop.
- Evaluation & Benchmarks — the evaluator layer in full: why RL evaluation is hard, the benchmark landscape, reporting standards.
- Offline Reinforcement Learning — the full expansion of the "offline" branch among the three data flows: the OOD problem and its solutions.
- Reward Engineering — the most underrated part of the environment layer: the reward is the product requirement.
- Choosing Frameworks & Tools — how to choose ready-made implementations for the learning-algorithm and experience-collection layers.
- Common Pitfalls & Anti-Patterns — a master table of the most common failure modes in each of the six layers.
References
- Farama Foundation. Gymnasium Documentation. https://gymnasium.farama.org/ — the official standard for the environment layer's interface contract (
reset/step/render, observation and action spaces). - Stable-Baselines3 Contributors. Stable-Baselines3 Documentation. https://stable-baselines3.readthedocs.io/ — production-ready, off-the-shelf implementations of the learning-algorithm layer.
- Ray Project. Ray RLlib Documentation. https://docs.ray.io/en/latest/rllib/index.html — the representative implementation of parallelization and distributed training for the experience-collection layer.
- Weng, L. (2018). The Challenges of Reinforcement Learning. https://lilianweng.github.io/posts/2018-02-19-rl-overview/ — a systematic discussion of "what makes RL engineering hard," echoing this page's "the environment is the product" argument.
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. https://arxiv.org/abs/1709.06560 — required reading for the evaluator layer: the statistical traps of reproducibility and fair comparison in RL experiments.
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.), §1.1–1.3. MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the authoritative source for the agent-environment interface concept.