Appearance
Offline Reinforcement Learning
In one line: this page explains offline reinforcement learning (offline RL) — learning a policy from a single fixed dataset, never interacting with the environment again. After reading it, you'll be able to explain the mechanism behind the core problem of "out-of-distribution actions + value overestimation," understand the three implementations of the conservatism idea (policy constraint, conservative value, implicit Q), and judge against real business scenarios whether offline RL is the right tool.
1. Problem Setup and Motivation: Why "No Environment Access" Is the Norm
1.1 The setup: from data to policy, with zero online interaction
Online RL follows a "learn while doing" paradigm — trial and error, exploration, watching the feedback. Offline RL is a completely different beast:
text
The offline RL paradigm
┌───────────────────────────────────────────────────────┐
│ 1. A fixed dataset D = {(s, a, r, s')} │
│ (collected by some behavior policy / expert / │
│ historical logs) │
│ 2. Training: learn only on D — interacting with the │
│ environment is forbidden │
│ 3. Deployment: the trained policy goes live, and only │
│ then meets the real world │
└───────────────────────────────────────────────────────┘Not a single environment interaction is allowed during training. Why is this setup so common?
| Motivation | Examples |
|---|---|
| Safety | Self-driving cars, medicine, nuclear reactors: online trial and error can kill |
| Cost | Real robot hardware, live financial trading: one trial can be extremely expensive |
| Compliance / ethics | You can't run wild experiments on real users with ads (see RL in Recommendation and Advertising) |
| Data already exists | Historical logs have already piled up to hundreds of millions of entries (the logs are the data) |
1.2 Don't confuse it with "experience replay" — they're not the same thing
The "offline" in offline RL is not the "off" in off-policy. Off-policy training still explores online (the behavior policy and the target policy differ, but both interact with the environment); offline RL forbids interaction entirely — it is, in effect, "hindsight learning on somebody else's data (or your own past data)." That's what makes it brutally hard and hugely valuable in engineering: the dataset is literally the only thing you have.
2. The Core Difficulty: OOD Actions and Value Overestimation (Bootstrap Hallucination)
2.1 The mechanism: the value function "gets high on its own supply" in unvisited places
Train a Q function on the dataset D — say, with the Q-learning objective from Value Learning:
$$ Q(s,a) \leftarrow r + \gamma \max_{a'} Q(s', a') $$
The root cause lies in $\max_{a'}$: that max can pick an action that never appears in the dataset — because Q is a neural network, it outputs a value for any input, even for an (s', a') pair that was never sampled.
When $Q(s', a')$ overestimates an OOD action (a neural network typically produces unreasonable outputs on inputs it has never seen — too high or too low), the max will lock onto that inflated action, and then Q trains itself on the "inflated Q" as its own target — a bootstrap hallucination: Q grows more and more inflated, until the entire network is overestimated and the values are completely distorted.
text
The bootstrap-hallucination vicious cycle
┌───────────────────────────────────────────────────────┐
│ Q gives an inflated value to an action it has │
│ never seen │
│ │ │
│ ▼ │
│ max_a' Q(s', a') picks the inflated action │
│ │ │
│ ▼ │
│ train Q on the inflated target ──▶ Q becomes even │
│ more inflated │
│ │ │
│ └── cycle until the value function │
│ blows up / distorts ────────────────────┘
└───────────────────────────────────────────────────────┘The essential difference from online RL
Online RL has OOD problems too, but the policy can act, explore, and correct its wrong estimates. Offline RL has no such escape hatch — the dataset is sealed, and Q's hallucinations about OOD actions never get corrected by reality. So the entire art of offline RL boils down to one principle: "don't let Q trust actions it has never seen."
2.2 A visual sketch
text
Q value
▲
│ ╱╲ ← true Q (unknowable)
│ ╱╲ ╱ ╲
│ ╱ ╲ ╱ ╲ ← learned Q (inflated where there's no data)
│ ╱ ╳ ╲
│ ╱ ╱ ╲ ╲
│ ╱ ╱ ╲ ╲
│ ╱ ╱ ╲ ╲
└─────────────╳───────────────► action a
│
the region covered by the data (dashed line);
everything beyond it is the "hallucination zone"2.3 Why the word "out-of-distribution" deserves your full attention
At deployment, the policy will visit state-action regions that the dataset covers only sparsely. Judging an offline policy means judging it over the entire state space, while the dataset only covers the places the behavior policy visited. Distribution mismatch → generalization failure. That's why offline RL papers keep repeating words like "support," "conservative," and "stay close to the data."
3. Three Families of Solutions: Policy Constraint / Conservative Value / Implicit Q
3.1 Family 1: policy constraint (BCQ, BEAR) — don't stray too far
The idea: constrain the policy to the "support region" of the dataset — the model may only pick actions it has "seen in the data."
- BCQ (Batch-Constrained Q-learning, Fujimoto et al., 2019): a conditional VAE (CVAE) models "the distribution of actions that appear at state s in the data"; the policy samples a handful of candidate actions from this distribution and picks the one with the highest Q. The action space is "clamped" to the data range, so OOD actions never even get a chance.
- BEAR (Bootstrapping Error Accumulation Reduction, Kumar et al., 2019): uses a "support distance" (MMD) to keep the policy close to the data distribution, allowing a more flexible family of distributions than BCQ.
A rough analogy: BCQ is "only order dishes from recipes you've seen before"; BEAR is "order from variants near those recipes."
3.2 Family 2: conservative value (CQL) — teach Q some humility
The idea: don't let Q overestimate — actively push down the Q of OOD actions. CQL (Conservative Q-Learning, Kumar et al., 2020) adds a penalty term on top of the standard Q-learning objective: make "Q on actions outside the data distribution" as small as possible.
$$ Q = \arg\min_Q \Big[ \underbrace{\alpha , \mathbb{E}{(s,a)\sim D}[Q(s,a)]}{\text{push down } Q \text{ on in-distribution actions}} - \alpha , \mathbb{E}{a'\sim \pi}[\dots] + \underbrace{\text{standard TD error}}{\text{keep the learning correct}} \Big] $$
(The full formula is rather heavy; the core intuition is: learn normally on the (s, a) pairs seen in the data, and push Q down on actions never seen, so that $\max_{a'} Q(s',a')$ no longer picks hallucinated actions.)
Where CQL wins: it bakes "conservatism" into the objective, works reliably, and has become one of the de facto baselines for offline RL. Variants include CQL(𝒟) and Cal-QL.
3.3 Family 3: implicit Q (IQL) — sidestep the max altogether
The idea: skip learning Q and skip the max — learn "expectiles" instead.
IQL (Implicit Q-Learning, Kostrikov et al., 2022): uses expectile regression (an asymmetric, "one-sided" regression) to learn the value function — it only fits "the part of the data that performs better than average":
$$ L(\theta) = \mathbb{E}\left[ L_2^\tau\left( r + \gamma V(s') - Q_\theta(s,a) \right) \right] $$
where $L_2^\tau(u) = |\tau - \mathbf{1}{u<0}| u^2$. The intuition: with τ > 0.5, the regression "leans toward higher values," approximating "the best value achievable at that state," yet it never needs a max over OOD actions — because its policy improves by sampling actions seen in the data.
IQL's core appeal: it doesn't require (state, optimal-action) pairs in the offline data, so it is more robust to suboptimal or noisy data; and it is simple to implement (nothing but expectile regression + behavior-cloning-style policy extraction).
3.4 The three families side by side
| Method | Idea | Data requirements | Implementation difficulty | Character |
|---|---|---|---|---|
| BCQ | Constrain the policy (restricted actions) | Sufficient data coverage | Medium | Conservative, stable |
| CQL | Conservative value (push down OOD Q) | Medium | Medium-high | Strong performance; a standard baseline |
| IQL | Implicit Q (sidestep the max) | Low (suboptimal data OK) | Low | Simple, robust |
4. The Relationship to Behavior Cloning / Imitation Learning
Offline RL's closest relative is behavior cloning (BC) — both learn from offline data only, and neither interacts online. The difference lies in what they learn:
| Dimension | Behavior cloning (BC) | Offline RL |
|---|---|---|
| Learning target | Imitate the behavior policy's (s→a) mapping | Optimize return (beat the behavior policy) |
| Labels required | Action labels | Rewards |
| Can it beat the data? | No (it can copy at best) | In principle yes (find a policy better than the data) |
| Failure mode | Distribution shift (compounding errors) | Value hallucination (OOD overestimation) |
| Data requirements | Expert data preferred | Good coverage + rewards |
The key insight
Offline RL's promise is "learning a policy better than the data from suboptimal data" (by judging value). Behavior cloning can't do that. But the converse also holds: when data quality is poor or rewards are unreliable, BC is the safer baseline — run BC first, then try offline RL, and use offline evaluation to decide which is stronger.
From imitation to RL: middle grounds like DAgger
DAgger (Dataset Aggregation) lets the policy explore during training and asks an expert to label the visited states — it needs an "expert" but never scores interactions with the environment. It is an interactive variant of imitation learning, sitting between BC and RL. For the related discussion, see RL vs. Neighboring Fields.
5. The Evaluation Trap: Why Offline Evaluation Is Hard
The most underrated difficulty of offline RL is evaluation: if interaction is forbidden during training, how do you know whether the new policy is any good? All you have is the dataset D.
5.1 The direct approaches — and their problems
- Compute return on the dataset: the (s,a) pairs the new policy visits may not be in the data, so their rewards can't simply be "looked up";
- Evaluate with the learned Q: Q itself hallucinates on OOD inputs, so the evaluation is untrustworthy — hallucination judging hallucination, a circular argument;
- Importance sampling (IS): correct in theory, but the variance explodes exponentially with the trajectory length.
5.2 Practical choices
| Method | What it does | Limitation |
|---|---|---|
| Trajectory-level IS | Re-weight the trajectories in the data | Variance explodes with length |
| Model-based evaluation | Learn a world model and simulate | Model error propagates through |
| Conservative estimation | Use a "no-overestimation" value like CQL/IQL as the evaluator | Gives a lower bound only, no exact number |
| Pre-launch A/B | Small-traffic canary with guardrails | Essentially "small-scale online" |
An iron rule of engineering
Offline evaluation can only give you a "relative ranking," never "absolute performance." The workflow that actually works in industry: rank a handful of candidate policies offline, then verify with a small-traffic canary. Treat any tool that claims to "predict online results precisely from offline evaluation" with suspicion.
6. Offline Data Engineering and the Deployment Pipeline
6.1 Data quality: the first productivity lever of offline RL
Everything in offline RL rests on the data — how good the data is sets how high the ceiling is. Three dimensions you must audit:
| Dimension | What to audit | Common problems |
|---|---|---|
| Coverage | Whether the state-action space is sufficiently visited | The behavior policy is too homogeneous → large regions have no data → OOD is unavoidable |
| Behavior-policy diversity | Whether the data comes from multiple policies / multiple periods | Single-policy data → the learned policy is just an "interpolation" of it |
| Reward consistency | Whether rewards are defined consistently and unpolluted | Historical logs change the reward definition midway → contradictory value signals |
Practical advice: run a "data profile" before training — statistics on the sparsity of the (s,a) distribution, the variance of returns in each region, and distribution drift over time. Data profiling should be as formal a step as model training, not an afterthought during debugging.
6.2 The deployment pipeline: offline for candidates, online for verification
The industry-standard pipeline (works for recommendations, advertising, and robotics logs alike):
text
Five steps to put offline RL into production
1. Data prep: cleaning, dedup, reward calibration, profiling
2. Multiple candidates: a BC baseline + 2-3 offline algorithms
(CQL/IQL/...) each produce a candidate policy
3. Offline ranking: rank the candidates with conservative
evaluation (CQL value / IS)
4. Online canary: small-traffic A/B with guardrails (limited
exploration, rollback capability)
5. Data feedback: canary-period data flows into the next round
of training (an online-offline loop)Why the canary step is non-negotiable
Offline evaluation can only rank, not pin down absolute numbers (see the previous section). Only a small-traffic canary can answer "does this policy actually work in the real world?" — and if it blows up, the damage is contained. The right way to run offline RL is as a fast-iterating online-verification pipeline, not as a one-click replacement.
6.3 Connecting to model-based: offline world models
A complementary line that has gained traction recently: learn a world model from offline data, then do more "offline exploration" inside it (offline model-based RL):
- The upside: the model can "imagine" trajectories the data never covered, easing the coverage problem;
- The risk: the imagined parts carry model error (see the "model-error trap" in Model-Based RL and World Models) — so evaluation inside the model must be conservative or uncertainty-weighted.
Methods of this family (e.g., MOReL, COMBO) are still a research frontier; see Frontier Progress.
7. Deployment Reality: Recommendations, Finance, Robotics
7.1 Recommender systems: the biggest stage for offline RL
Recommendation naturally satisfies every precondition of offline RL: massive historical behavior logs, expensive online trial and error, and ready-made rewards (clicks/conversions). The deployment mix in practice:
- Short term: train the policy offline on log data, and keep collecting data after it goes live;
- Conservatism: because user behavior is non-stationary (today's data expires tomorrow), industry practice is usually a hybrid recipe — offline RL for the policy + small-traffic online verification + periodic retraining. See RL in Recommendation and Advertising.
7.2 Financial trading
Financial data has an extremely low signal-to-noise ratio, and markets are non-stationary — offline RL's "data as history" assumption barely holds there. The realistic conclusion: in finance, offline RL fits quasi-stationary sub-problems like execution optimization far better than whole trading strategies (for the full analysis, see RL in Financial Trading).
7.3 Robotics: learning from logs
Robotics labs sit on mountains of "human demonstrations + old-policy logs" (teleoperation data, multi-round experiment logs). Offline RL lets a new policy learn from this historical data without fresh online runs every time — the core application direction in robotics. See Robotics Control and Sim2Real and Frontier Progress.
8. Decision Tree: When Offline RL Is the Right Call
text
You have a historical dataset + a sequential decision problem
│
├── Were the actions in the data collected by a behavior
│ policy, with good coverage?
│ ├── No ──▶ don't use offline RL (the data isn't up to it;
│ │ start with BC / go online instead)
│ └── Yes ↓
├── Is the reward signal reliable?
│ ├── No ──▶ build a reward first with IRL / preference
│ │ learning (see the reward engineering page)
│ └── Yes ↓
├── Can you safely validate online at small scale?
│ ├── Yes ──▶ offline RL for candidates + small-traffic A/B
│ │ (recommended)
│ └── No ──▶ proceed with caution: offline evaluation can
│ only rank — don't bet on absolute performance
│
Bottom line: offline RL fits "data already available + reliable
rewards + canary validation possible"; it does not fit "bad data +
fabricated rewards + one-click full rollout."An engineering checklist
- Build the BC baseline first: BC typically reaches 70-90% of the behavior policy's level — the floor that offline RL must beat;
- Audit the data: check data coverage (is the behavior policy too homogeneous?), reward consistency, and contamination from environment bugs (see Common Pitfalls and Anti-Patterns);
- Start with the conservatism knobs turned up: tune CQL's α, BCQ's number of candidates, and similar "conservative dials" toward the safe side first, and relax only once things are stable;
- Rank offline, verify with a canary: never forget the iron rule from the previous section.
Frontier directions
The intersection of offline RL with "world models" and "diffusion-model policies" is a current hotspot (e.g., Diffuser, IQL + diffusion). For newer developments — the RvS debate, model-based offline RL — see Frontier Progress.
Further Reading
- RL in Recommendation and Advertising — offline RL's most mature landing zone: log data + canary verification
- RL in Financial Trading — offline RL's traps in non-stationary markets, and "when it actually works"
- Value Learning: From Dynamic Programming to DQN — the max operator in Q-learning is the source of OOD hallucinations
- Policy Gradient Methods — how offline policy gradient (offline PG) handles things with importance weighting
- Frontier Progress — the latest offline RL work beyond IQL/CQL
- Common Pitfalls and Anti-Patterns — contaminated data, evaluation cheating, and other frequent offline-project failures
References
- Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. arXiv:2005.01643 (the offline RL survey — required reading)
- Fujimoto, S., Meger, D., & Precup, D. (2019). Off-Policy Deep Reinforcement Learning without Exploration. ICML. arXiv:1812.02900 (BCQ)
- Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q-Learning for Offline Reinforcement Learning. NeurIPS. arXiv:2006.04779 (CQL)
- Kostrikov, I., Nair, A., & Levine, S. (2022). Offline Reinforcement Learning with Implicit Q-Learning. ICLR. arXiv:2110.06169 (IQL)
- Kumar, A., Fu, J., Soh, M., Tucker, G., & Levine, S. (2019). Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. NeurIPS. arXiv:1906.00949 (BEAR)
- Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2020). D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv:2004.07219 (the D4RL benchmark for offline RL)