Appearance
RL vs Neighboring Paradigms
One-line pitch: this page answers a practical question — "does my problem deserve RL?" It lines RL up against six neighboring paradigms (supervised learning, unsupervised learning, optimal control, operations research, planning & search, behavior cloning), compares them one pair at a time, and ends with an actionable selection decision table. When you're done, you should be able to say one of "use RL," "don't use RL," or "start with something else and transition to RL later" about your own task — instead of the wishy-washy "let's give RL a try."
1. The Verdict First: Six Paradigms, Each Missing One Dimension
RL is not a lonely discipline — it stands in a crowd of neighbors. Here's the bird's-eye view first:
| Paradigm | One-liner | Key difference from RL | Relationship |
|---|---|---|---|
| Supervised learning | Learn a mapping from (x, y) | Data is stationary, feedback is immediate, no sequential decisions | Supplies function approximation and data thinking |
| Unsupervised learning | Find structure in data | No notion of "decision → consequence" | Usable for state representation learning |
| Optimal control | With a known model, find the optimal control sequence | Usually assumes a known model; goal is trajectory optimization | The "learning-free" version of RL when the model is known |
| Operations research | Find the optimal solution to a static problem | Usually no time dimension, no interactive learning | A sequential OR problem is an MDP |
| Planning & search | With a known model, find a path by search | No "learning from data"; computationally heavy | Can combine with learning (MCTS + network) |
| Behavior cloning | Copy the expert's demonstrations | No trial and error, no exploration, capped by the expert | RL's "cold start" tool |
Memorize it in one line: RL = sequential decision-making + trial-and-error learning + long-term return maximization; every neighbor is missing at least one of the three. Now let's break the pairs down one by one.
2. Versus Supervised and Unsupervised Learning
This part already got a full three-dimension treatment (data form / feedback / goal) in What Is Reinforcement Learning, so no need to repeat the table here. Instead, three engineering differences that are easy to overlook:
- Data is not independent and identically distributed: supervised learning assumes independent samples; in RL, your next state is caused by your previous action — which directly motivates engineering tricks like experience replay that "forcibly break the correlation." See Value Learning.
- Training and deployment are inseparable: supervised learning trains offline and then ships; for RL, deployment (interacting with the environment) is part of training — hence the entire subfield of Offline RL, dedicated to "training without further interaction."
- There is no fixed "test set": RL evaluation must cope with environment stochasticity, seed sensitivity, and non-stationarity — an order of magnitude harder. See Evaluation & Benchmarks.
Unsupervised learning's hidden role
Unsupervised learning isn't RL's rival; it's a component: world models learn "low-dimensional representations of states" (latent spaces), and count-based intrinsic rewards use density estimation. To see this convergence in action, visit Model-Based RL and World Models.
3. RL vs Optimal Control
Optimal control is RL's close cousin under the "known model" assumption. Both solve sequential decision-making over continuous state spaces, and both are mathematically related (they lead to the same place: dynamic programming and the Bellman equation). The dividing line is whether the model is known:
| Dimension | Optimal control | RL | Notes |
|---|---|---|---|
| Dynamics model | Usually known (or precisely identifiable) | Usually unknown, learned from samples | This is the most fundamental dividing line |
| Objective | Minimize a cost function J(u) | Maximize cumulative reward G | Two notations for the same problem |
| Toolbox | LQR, calculus of variations, PMP, MPC | TD, policy gradients, Q-learning | MPC is "receding horizon + replanning" |
| Handling randomness | Classical methods mostly assume deterministic/Gaussian systems | Handles stochastic transitions and stochastic policies natively | — |
| State dimension | Often low-dimensional continuous (robotics, aerospace) | Can be high-dimensional discrete/continuous (images, text) | — |
MPC (model predictive control) deserves a closer look: at every small step, it rolls a known model forward to optimize a few steps, executes the first action, then replans. Its relationship to model-based RL:
- MPC doesn't learn — it only optimizes; feed it a wrong model and it breaks;
- model-based RL fills in the "learn the model" step: first learn a world model, then solve within it using MPC or planning — see Model-Based RL;
- modern work such as Dreamer and TD-MPC welds the two together: learn a latent-space model + MPC-style planning.
Engineering reality
Real physical systems (robots, aircraft) often have precise analytical models (rigid-body dynamics), and here optimal control is usually more stable, more interpretable, and more sample-efficient than RL. RL's stage is the scenarios where models are too hard to build (soft bodies, fluids, contact-rich manipulation). That's why in industry, "LQR/MPC first; RL only for the parts you can't write a model for." For a reality check from the robotics side, see Robotics Control and Sim2Real.
4. RL vs Operations Research (OR)
Operations research deals with "finding the optimal solution under constraints": linear programming, integer programming (MILP), combinatorial optimization (TSP, bin packing, scheduling). Its relationship with RL is often misunderstood, so it's worth splitting into two levels.
1. Static OR is not RL — but "sequentialized," it becomes RL
| Aspect | Classic OR form | Once written as an MDP |
|---|---|---|
| Time | A single decision (solve one optimization problem) | Step-by-step decisions (each step solves a "sub-problem") |
| State | None / static | Remaining tasks, inventory, machine occupancy |
| Action | One-shot assignment/scheduling | Assignment/scheduling decision at each step |
| Objective | Minimize total cost | Minimize long-term discounted cost |
Inventory management is the textbook example: the static newsvendor problem (how much to order once) is OR; dynamic inventory control (decide every period how much to order, with next period's inventory becoming the state) is a standard MDP — Bellman originally invented dynamic programming for exactly this. The same goes for scheduling, bin packing, and network routing — the test is "does this decision change the board for the next decision?" To dig deeper, see RL in Scheduling and Operations Research.
2. Learning-based solvers: RL as a "construction heuristic"
Classic OR solvers (MILP, branch and bound) explode exponentially as problem size grows. The new use of RL is to train a policy that constructs or improves solutions directly (learning to search / learning to construct):
text
Classic route: exact solver (MILP) ──→ times out once size grows
Heuristic route: hand-crafted rules ──→ fast but suboptimal
Learning route: RL-trained construction policy ──→ fast + adapts to problem distributionThe value of RL here isn't "replacing MILP" — it's delivering speed in online scenarios (routing, scheduling) where every decision needs an approximate answer in real time. The evaluation difficulty lies in generalization: whether the training problem distribution covers the real one. See Evaluation & Benchmarks.
A shared ancestry
OR and RL have the same founding father — Bellman. Dynamic programming is used by the OR community as an optimization tool, and it is also RL's mathematical foundation (the Bellman equation on the Markov Decision Process page is his invention). The split between the two fields is mostly a difference of research communities, not of mathematics.
5. RL vs Planning & Search
Classical AI planning (the STRIPS/PDDL lineage) and search (A*, MCTS) solve the same kind of problem — "find a sequence of actions that reaches a goal." The dividing line with RL, again, is: is the model known, and do you need to learn?
| Dimension | Classical planning/search | RL |
|---|---|---|
| Model | Given explicitly (action effects, goal) | Unknown by default, learned from interaction |
| How action sequences are found | Search / backtracking / heuristics | Gradient-based learning of policy/value functions |
| Generalization to new instances | Poor (search from scratch each time) | Good (learns the "knack") |
| Where the computation happens | At decision time, in the search | In the training phase |
MCTS (Monte Carlo tree search) is the most beautiful meeting point on this boundary: it is itself "efficient search with a known model," but AlphaGo combined it with policy and value networks — the networks supply "priors" and "evaluations" to the search, and the search spends its compute on the branches that matter. That's the "search × learning" principle behind AlphaGo and Monte Carlo Tree Search, and the real-world form of "planning in your head" in Model-Based RL.
An engineering heuristic
If your domain has a precise, callable simulator (say, the rules of Go, or a circuit simulator), search/planning often gives you a strong solution immediately — no RL needed. Only when the simulator is too expensive (can't afford deep search at every step) or the model is inaccurate do you need "learning to replace search," or "learning + shallow search." AlphaGo's victory came precisely from the latter: search depth is limited, so let the network "guess" the parts that went unsearched.
6. RL vs Behavior Cloning / Imitation Learning
Behavior cloning (BC), the most straightforward branch of imitation learning, treats expert demonstrations as supervised learning data and learns π(a|s). It's the neighbor most easily confused with RL — and the one most often misused — because it's so easy to implement:
python
# behavior cloning: it's just multi-class classification / regression!
# data: expert trajectories [(s1,a1), (s2,a2), ...]
# objective: minimize cross-entropy between π(a|s) and the demonstrated action
loss = cross_entropy(policy(s), expert_action)1. BC's three fatal flaws
| Limitation | Cause | Consequence |
|---|---|---|
| Distribution shift (compounding error) | Training only ever sees expert states; once you drift off the expert trajectory at test time, errors compound | Long-horizon tasks diverge from the expert exponentially |
| The ceiling is the expert | What you learn is capped at the demonstrator's level | Can never surpass the demonstration (superhuman Go would be out of the question) |
| No feedback about mistakes | Never learns "doing this earns a negative reward" | Helpless in situations outside the demonstrations |
2. Three variants of imitation learning
| Variant | Idea | Representative work | Relationship to RL |
|---|---|---|---|
| Behavior cloning (BC) | Directly regress the expert's actions | Pomerleau 1989 ALVINN | No RL; pure supervised learning |
| Inverse RL (IRL) | Infer the reward function from demonstrations | Ng & Russell 2000; MaxEnt IRL | The reward comes from demonstrations, but RL is still needed afterwards to solve it |
| Generative adversarial imitation (GAIL) | A discriminator tells "my trajectory vs the expert's" apart; the policy learns to fool it | Ho & Ermon 2016 | Formally resembles adversarial training; needs no explicit reward |
Imitation learning's real-world role is RL's cold starter: first train a decent policy with BC, then fine-tune with RL to surpass the expert. Autonomous driving is the classic example of this path — copy first, learn later; see Autonomous Driving Decision-Making. In the LLM context, SFT is "behavior cloning," and RLHF is "the RL fine-tuning after behavior cloning" — a relationship hammered home again and again in RLHF and Alignment with Human Feedback.
The most common misjudgment
Reading "we have lots of human demonstrations" as "we should use RL." If all you have is demonstrations — no environment, no reward — behavior cloning alone buys you 80% of the value, and RL is the harder path. The right order is: BC baseline first → environment available for trial and error → then bring in RL. Do it backwards — jumping straight to RL — and the exploration space is so vast that learning usually goes nowhere.
7. Where Deep Learning Sits in RL (Deep RL)
Deep RL = deep learning provides the function approximation + RL provides the learning objective. Broken down:
| Deep learning's role | Where it sits in RL | Corresponding page |
|---|---|---|
| Value function approximation | DQN uses a CNN to estimate Q(s,a) from pixels | Value Learning |
| Policy function approximation | A policy network outputs the action distribution π(a|s) | Policy Gradient Methods |
| State representation/compression | World models encode high-dimensional observations into a latent space | Model-Based RL |
| Reward/value modeling | Reward models (Bradley–Terry scorers) | RLHF |
So "deep RL" is not a new paradigm — it's the same paradigm with upgraded representational power. The price is equally clear: deep learning brings tuning sensitivity, low sample efficiency, and poor interpretability; the engineering responses live in Tuning in Practice and Common Pitfalls.
Deep learning in one sentence
"Deep learning decides how complex a pattern the agent can memorize; RL decides why it should memorize it." The former answers representational capacity; the latter answers the learning objective — you need both.
8. The Selection Decision Table
Here's a decision table condensing all six comparisons. Ask yourself four questions, in order, and the answers walk you to a verdict:
| Question | Yes | No |
|---|---|---|
| Q1. Is this sequential decision-making (actions affect future situations)? | Go to Q2 | Plain supervised/unsupervised/static optimization is enough — no RL needed |
| Q2. Do you have a precise, usable model or simulator? | Prefer planning/search or optimal control; consider model-based RL only if simulation is expensive | Go to Q3 |
| Q3. Do you have plenty of expert demonstrations? | Start with behavior cloning as a baseline, then transition to RL | Go to Q4 |
| Q4. Can you write a trustworthy reward and afford the cost of trial and error? | ✅ Go with RL (start by Building an RL Project from Scratch) | ❌ Consider Offline RL, or don't do it at all |
A second, finer-grained table for specific situations:
| Specific situation | Recommended paradigm | Rationale |
|---|---|---|
| Robot trajectory tracking; the model can be precisely identified | Optimal control (MPC/LQR) | Stable, fast, interpretable — RL is unnecessary |
| Games / model-free games and adversarial play | RL (+ MCTS hybrid) | Exploration is the only way through |
| Recommendation, ads, A/B testing (single-step choice) | Contextual bandit | Simpler than full RL and good enough; see Multi-Armed Bandits |
| Inventory, scheduling, bin packing (long-term costs) | Sequential decision-making → RL or dynamic programming | At large scales, learn a construction heuristic |
| Plenty of expert logs; online trial and error impossible | Behavior cloning + offline RL | See Offline RL |
| Goal hard to write as a scalar reward | IRL/preference learning first, or re-evaluate whether the goal can be simplified | See Reward Engineering |
| Want to accelerate LLM inference/alignment | RLHF / direct preference optimization | See RLHF |
One final word of advice
The cost of a wrong choice lies not in "chose RL," but in choosing RL without its three-piece kit — environment, reward, and evaluation — in place. Among RL's failure stories, the vast majority aren't "the algorithm didn't work" but "the problem wasn't a fit + the infrastructure didn't keep up." After making your decision, read the pitfall summary in Common Pitfalls and Anti-Patterns once more and check your project against it.
Further Reading
- What Is Reinforcement Learning — the "baseline plane" for this page's comparison tables: RL's own definition, four elements, and minimal example.
- Markov Decision Process — the strict mathematical criterion for "when can a problem be written as RL."
- Model-Based RL and World Models — the full development of the line where RL converges with optimal control and MPC.
- RL in Scheduling and Operations Research — real cases of static OR problems being sequentialized and handed to RL.
- Autonomous Driving Decision-Making — the engineering reality of the imitation learning + RL + safety constraints hybrid route.
- Robotics Control and Sim2Real — how optimal control and RL divide the labor in robotics today.
References
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.), §1.7 "Relationship to Optimal Control". MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the textbook's authoritative treatment of the boundaries among RL, optimal control, and operations research.
- Ho, J. & Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS 2016. https://arxiv.org/abs/1606.03476 — the GAIL paper, a key reference for the relationship between imitation learning and RL.
- Ng, A. Y. & Russell, S. (2000). Algorithms for Inverse Reinforcement Learning. ICML 2000. https://ai.stanford.edu/~ang/papers/icml00-irl.pdf — the pioneering paper on inverse reinforcement learning.
- Pomerleau, D. A. (1989). ALVINN: An Autonomous Land Vehicle in a Neural Network. NeurIPS 1989. — the classic origin of behavior cloning (an autonomous-driving prototype vehicle).
- Browne, C. et al. (2012). A Survey of Monte Carlo Tree Search Methods. IEEE Transactions on Computational Intelligence and AI in Games. https://ieeexplore.ieee.org/document/6145622 — a survey of MCTS, the background reading for combining planning/search with learning.
- OpenAI (2018). Spinning Up in Deep RL — Key Concepts. https://spinningup.openai.com/en/latest/spinningup/rl_intro.html — a concise English-language contrast of "why RL differs from supervised/unsupervised learning."