Skip to content

RL vs Neighboring Paradigms

On this page Boundary analysis of reinforcement learning versus supervised learning, unsupervised learning, optimal control, operations research, planning & search, and behavior cloning — a decision framework for "when to use RL, and when not to."

RL vs Neighboring Paradigms ​

One-line pitch: this page answers a practical question — "does my problem deserve RL?" It lines RL up against six neighboring paradigms (supervised learning, unsupervised learning, optimal control, operations research, planning & search, behavior cloning), compares them one pair at a time, and ends with an actionable selection decision table. When you're done, you should be able to say one of "use RL," "don't use RL," or "start with something else and transition to RL later" about your own task — instead of the wishy-washy "let's give RL a try."

1. The Verdict First: Six Paradigms, Each Missing One Dimension ​

RL is not a lonely discipline — it stands in a crowd of neighbors. Here's the bird's-eye view first:

ParadigmOne-linerKey difference from RLRelationship
Supervised learningLearn a mapping from (x, y)Data is stationary, feedback is immediate, no sequential decisionsSupplies function approximation and data thinking
Unsupervised learningFind structure in dataNo notion of "decision → consequence"Usable for state representation learning
Optimal controlWith a known model, find the optimal control sequenceUsually assumes a known model; goal is trajectory optimizationThe "learning-free" version of RL when the model is known
Operations researchFind the optimal solution to a static problemUsually no time dimension, no interactive learningA sequential OR problem is an MDP
Planning & searchWith a known model, find a path by searchNo "learning from data"; computationally heavyCan combine with learning (MCTS + network)
Behavior cloningCopy the expert's demonstrationsNo trial and error, no exploration, capped by the expertRL's "cold start" tool

Memorize it in one line: RL = sequential decision-making + trial-and-error learning + long-term return maximization; every neighbor is missing at least one of the three. Now let's break the pairs down one by one.

2. Versus Supervised and Unsupervised Learning ​

This part already got a full three-dimension treatment (data form / feedback / goal) in What Is Reinforcement Learning, so no need to repeat the table here. Instead, three engineering differences that are easy to overlook:

  1. Data is not independent and identically distributed: supervised learning assumes independent samples; in RL, your next state is caused by your previous action — which directly motivates engineering tricks like experience replay that "forcibly break the correlation." See Value Learning.
  2. Training and deployment are inseparable: supervised learning trains offline and then ships; for RL, deployment (interacting with the environment) is part of training — hence the entire subfield of Offline RL, dedicated to "training without further interaction."
  3. There is no fixed "test set": RL evaluation must cope with environment stochasticity, seed sensitivity, and non-stationarity — an order of magnitude harder. See Evaluation & Benchmarks.

Unsupervised learning's hidden role

Unsupervised learning isn't RL's rival; it's a component: world models learn "low-dimensional representations of states" (latent spaces), and count-based intrinsic rewards use density estimation. To see this convergence in action, visit Model-Based RL and World Models.

3. RL vs Optimal Control ​

Optimal control is RL's close cousin under the "known model" assumption. Both solve sequential decision-making over continuous state spaces, and both are mathematically related (they lead to the same place: dynamic programming and the Bellman equation). The dividing line is whether the model is known:

DimensionOptimal controlRLNotes
Dynamics modelUsually known (or precisely identifiable)Usually unknown, learned from samplesThis is the most fundamental dividing line
ObjectiveMinimize a cost function J(u)Maximize cumulative reward GTwo notations for the same problem
ToolboxLQR, calculus of variations, PMP, MPCTD, policy gradients, Q-learningMPC is "receding horizon + replanning"
Handling randomnessClassical methods mostly assume deterministic/Gaussian systemsHandles stochastic transitions and stochastic policies natively—
State dimensionOften low-dimensional continuous (robotics, aerospace)Can be high-dimensional discrete/continuous (images, text)—

MPC (model predictive control) deserves a closer look: at every small step, it rolls a known model forward to optimize a few steps, executes the first action, then replans. Its relationship to model-based RL:

  • MPC doesn't learn — it only optimizes; feed it a wrong model and it breaks;
  • model-based RL fills in the "learn the model" step: first learn a world model, then solve within it using MPC or planning — see Model-Based RL;
  • modern work such as Dreamer and TD-MPC welds the two together: learn a latent-space model + MPC-style planning.

Engineering reality

Real physical systems (robots, aircraft) often have precise analytical models (rigid-body dynamics), and here optimal control is usually more stable, more interpretable, and more sample-efficient than RL. RL's stage is the scenarios where models are too hard to build (soft bodies, fluids, contact-rich manipulation). That's why in industry, "LQR/MPC first; RL only for the parts you can't write a model for." For a reality check from the robotics side, see Robotics Control and Sim2Real.

4. RL vs Operations Research (OR) ​

Operations research deals with "finding the optimal solution under constraints": linear programming, integer programming (MILP), combinatorial optimization (TSP, bin packing, scheduling). Its relationship with RL is often misunderstood, so it's worth splitting into two levels.

1. Static OR is not RL — but "sequentialized," it becomes RL ​

AspectClassic OR formOnce written as an MDP
TimeA single decision (solve one optimization problem)Step-by-step decisions (each step solves a "sub-problem")
StateNone / staticRemaining tasks, inventory, machine occupancy
ActionOne-shot assignment/schedulingAssignment/scheduling decision at each step
ObjectiveMinimize total costMinimize long-term discounted cost

Inventory management is the textbook example: the static newsvendor problem (how much to order once) is OR; dynamic inventory control (decide every period how much to order, with next period's inventory becoming the state) is a standard MDP — Bellman originally invented dynamic programming for exactly this. The same goes for scheduling, bin packing, and network routing — the test is "does this decision change the board for the next decision?" To dig deeper, see RL in Scheduling and Operations Research.

2. Learning-based solvers: RL as a "construction heuristic" ​

Classic OR solvers (MILP, branch and bound) explode exponentially as problem size grows. The new use of RL is to train a policy that constructs or improves solutions directly (learning to search / learning to construct):

text
Classic route: exact solver (MILP)  ──→ times out once size grows
Heuristic route: hand-crafted rules ──→ fast but suboptimal
Learning route: RL-trained construction policy ──→ fast + adapts to problem distribution

The value of RL here isn't "replacing MILP" — it's delivering speed in online scenarios (routing, scheduling) where every decision needs an approximate answer in real time. The evaluation difficulty lies in generalization: whether the training problem distribution covers the real one. See Evaluation & Benchmarks.

A shared ancestry

OR and RL have the same founding father — Bellman. Dynamic programming is used by the OR community as an optimization tool, and it is also RL's mathematical foundation (the Bellman equation on the Markov Decision Process page is his invention). The split between the two fields is mostly a difference of research communities, not of mathematics.

Classical AI planning (the STRIPS/PDDL lineage) and search (A*, MCTS) solve the same kind of problem — "find a sequence of actions that reaches a goal." The dividing line with RL, again, is: is the model known, and do you need to learn?

DimensionClassical planning/searchRL
ModelGiven explicitly (action effects, goal)Unknown by default, learned from interaction
How action sequences are foundSearch / backtracking / heuristicsGradient-based learning of policy/value functions
Generalization to new instancesPoor (search from scratch each time)Good (learns the "knack")
Where the computation happensAt decision time, in the searchIn the training phase

MCTS (Monte Carlo tree search) is the most beautiful meeting point on this boundary: it is itself "efficient search with a known model," but AlphaGo combined it with policy and value networks — the networks supply "priors" and "evaluations" to the search, and the search spends its compute on the branches that matter. That's the "search × learning" principle behind AlphaGo and Monte Carlo Tree Search, and the real-world form of "planning in your head" in Model-Based RL.

An engineering heuristic

If your domain has a precise, callable simulator (say, the rules of Go, or a circuit simulator), search/planning often gives you a strong solution immediately — no RL needed. Only when the simulator is too expensive (can't afford deep search at every step) or the model is inaccurate do you need "learning to replace search," or "learning + shallow search." AlphaGo's victory came precisely from the latter: search depth is limited, so let the network "guess" the parts that went unsearched.

6. RL vs Behavior Cloning / Imitation Learning ​

Behavior cloning (BC), the most straightforward branch of imitation learning, treats expert demonstrations as supervised learning data and learns π(a|s). It's the neighbor most easily confused with RL — and the one most often misused — because it's so easy to implement:

python
# behavior cloning: it's just multi-class classification / regression!
# data: expert trajectories [(s1,a1), (s2,a2), ...]
# objective: minimize cross-entropy between π(a|s) and the demonstrated action
loss = cross_entropy(policy(s), expert_action)

1. BC's three fatal flaws ​

LimitationCauseConsequence
Distribution shift (compounding error)Training only ever sees expert states; once you drift off the expert trajectory at test time, errors compoundLong-horizon tasks diverge from the expert exponentially
The ceiling is the expertWhat you learn is capped at the demonstrator's levelCan never surpass the demonstration (superhuman Go would be out of the question)
No feedback about mistakesNever learns "doing this earns a negative reward"Helpless in situations outside the demonstrations

2. Three variants of imitation learning ​

VariantIdeaRepresentative workRelationship to RL
Behavior cloning (BC)Directly regress the expert's actionsPomerleau 1989 ALVINNNo RL; pure supervised learning
Inverse RL (IRL)Infer the reward function from demonstrationsNg & Russell 2000; MaxEnt IRLThe reward comes from demonstrations, but RL is still needed afterwards to solve it
Generative adversarial imitation (GAIL)A discriminator tells "my trajectory vs the expert's" apart; the policy learns to fool itHo & Ermon 2016Formally resembles adversarial training; needs no explicit reward

Imitation learning's real-world role is RL's cold starter: first train a decent policy with BC, then fine-tune with RL to surpass the expert. Autonomous driving is the classic example of this path — copy first, learn later; see Autonomous Driving Decision-Making. In the LLM context, SFT is "behavior cloning," and RLHF is "the RL fine-tuning after behavior cloning" — a relationship hammered home again and again in RLHF and Alignment with Human Feedback.

The most common misjudgment

Reading "we have lots of human demonstrations" as "we should use RL." If all you have is demonstrations — no environment, no reward — behavior cloning alone buys you 80% of the value, and RL is the harder path. The right order is: BC baseline first → environment available for trial and error → then bring in RL. Do it backwards — jumping straight to RL — and the exploration space is so vast that learning usually goes nowhere.

7. Where Deep Learning Sits in RL (Deep RL) ​

Deep RL = deep learning provides the function approximation + RL provides the learning objective. Broken down:

Deep learning's roleWhere it sits in RLCorresponding page
Value function approximationDQN uses a CNN to estimate Q(s,a) from pixelsValue Learning
Policy function approximationA policy network outputs the action distribution π(a|s)Policy Gradient Methods
State representation/compressionWorld models encode high-dimensional observations into a latent spaceModel-Based RL
Reward/value modelingReward models (Bradley–Terry scorers)RLHF

So "deep RL" is not a new paradigm — it's the same paradigm with upgraded representational power. The price is equally clear: deep learning brings tuning sensitivity, low sample efficiency, and poor interpretability; the engineering responses live in Tuning in Practice and Common Pitfalls.

Deep learning in one sentence

"Deep learning decides how complex a pattern the agent can memorize; RL decides why it should memorize it." The former answers representational capacity; the latter answers the learning objective — you need both.

8. The Selection Decision Table ​

Here's a decision table condensing all six comparisons. Ask yourself four questions, in order, and the answers walk you to a verdict:

QuestionYesNo
Q1. Is this sequential decision-making (actions affect future situations)?Go to Q2Plain supervised/unsupervised/static optimization is enough — no RL needed
Q2. Do you have a precise, usable model or simulator?Prefer planning/search or optimal control; consider model-based RL only if simulation is expensiveGo to Q3
Q3. Do you have plenty of expert demonstrations?Start with behavior cloning as a baseline, then transition to RLGo to Q4
Q4. Can you write a trustworthy reward and afford the cost of trial and error?✅ Go with RL (start by Building an RL Project from Scratch)❌ Consider Offline RL, or don't do it at all

A second, finer-grained table for specific situations:

Specific situationRecommended paradigmRationale
Robot trajectory tracking; the model can be precisely identifiedOptimal control (MPC/LQR)Stable, fast, interpretable — RL is unnecessary
Games / model-free games and adversarial playRL (+ MCTS hybrid)Exploration is the only way through
Recommendation, ads, A/B testing (single-step choice)Contextual banditSimpler than full RL and good enough; see Multi-Armed Bandits
Inventory, scheduling, bin packing (long-term costs)Sequential decision-making → RL or dynamic programmingAt large scales, learn a construction heuristic
Plenty of expert logs; online trial and error impossibleBehavior cloning + offline RLSee Offline RL
Goal hard to write as a scalar rewardIRL/preference learning first, or re-evaluate whether the goal can be simplifiedSee Reward Engineering
Want to accelerate LLM inference/alignmentRLHF / direct preference optimizationSee RLHF

One final word of advice

The cost of a wrong choice lies not in "chose RL," but in choosing RL without its three-piece kit — environment, reward, and evaluation — in place. Among RL's failure stories, the vast majority aren't "the algorithm didn't work" but "the problem wasn't a fit + the infrastructure didn't keep up." After making your decision, read the pitfall summary in Common Pitfalls and Anti-Patterns once more and check your project against it.

Further Reading ​

References ​

  • Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.), §1.7 "Relationship to Optimal Control". MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the textbook's authoritative treatment of the boundaries among RL, optimal control, and operations research.
  • Ho, J. & Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS 2016. https://arxiv.org/abs/1606.03476 — the GAIL paper, a key reference for the relationship between imitation learning and RL.
  • Ng, A. Y. & Russell, S. (2000). Algorithms for Inverse Reinforcement Learning. ICML 2000. https://ai.stanford.edu/~ang/papers/icml00-irl.pdf — the pioneering paper on inverse reinforcement learning.
  • Pomerleau, D. A. (1989). ALVINN: An Autonomous Land Vehicle in a Neural Network. NeurIPS 1989. — the classic origin of behavior cloning (an autonomous-driving prototype vehicle).
  • Browne, C. et al. (2012). A Survey of Monte Carlo Tree Search Methods. IEEE Transactions on Computational Intelligence and AI in Games. https://ieeexplore.ieee.org/document/6145622 — a survey of MCTS, the background reading for combining planning/search with learning.
  • OpenAI (2018). Spinning Up in Deep RL — Key Concepts. https://spinningup.openai.com/en/latest/spinningup/rl_intro.html — a concise English-language contrast of "why RL differs from supervised/unsupervised learning."