Skip to content

Model-Based RL and World Models

On this page Learn the environment's dynamics first, then plan "in your head" — MPC, Dyna, Dreamer, World Models, TD-MPC; the trade-off between data efficiency and model error; a model-based view of AlphaGo/AlphaZero.

Model-Based RL and World Models ​

In one line: this page explains model-based reinforcement learning (model-based RL) — instead of learning a policy directly, you first learn the "dynamics of the environment" (a world model) and then plan your actions inside that mental simulation. After reading it, you'll be able to state the essential difference between model-free and model-based, trace the three generations of techniques from MPC through Dyna to Dreamer, recognize the "model-error trap," and reread "learning × search" through AlphaGo's eyes.

1. Model-Free vs. Model-Based: Two Fundamentally Different Routes ​

1.1 The core difference: whether you "understand the world" ​

  • Model-free: never learns the world's regularities; it improves the value function or policy directly from trial-and-error samples. It answers only "what should I do?" (which action to pick), never "how does the world work?"
  • Model-based: first learns "if I take action a, what will the world look like and what reward will I get?" (transition + reward), and then plans inside the learned world model. It answers one extra question: "what happens if I do this?"
text
Model-free:  samples ──▶ value/policy (skips "understanding")
             only learns "what to do"

Model-based: samples ──▶ world model (ŝ, r̂) ──▶ mental planning ──▶ action
             first learn "how the world works", then roll it forward

1.2 Side by side ​

DimensionModel-freeModel-based
What it learnsQ(s,a) / π(as)
Sample efficiencyLow (millions of steps and up)High (tens of thousands of steps or fewer)
Where the final policy comes fromDirect optimizationPlanned from, or distilled out of, the model
Generalization to new tasksPoor (learn it all again)Good (the model is reusable and re-plannable)
Main riskBurning through samplesModel error (imagination bias)
RepresentativesDQN, PPO, SACDyna, Dreamer, TD-MPC, MuZero

The one-line intuition

Model-free is like "grinding practice problems until test-taking instinct kicks in"; model-based is like "learning the physics first and then solving the problems." The former burns a lot of samples but is blunt and simple; the latter saves samples, but "if you learned the physics wrong, everything downstream is wrong."

2. Learning a World Model: Transition + Reward ​

2.1 Formalization: the forward model ​

A world model predicts the outcome of a given state and action:

$$ \hat s_{t+1} = f_\phi(s_t, a_t), \qquad \hat r_{t+1} = g_\psi(s_t, a_t) $$

The training objective is plain supervised learning: on real interaction data $(s_t, a_t, s_{t+1}, r_{t+1})$, minimize the prediction error.

$$ \mathcal{L} = \mathbb{E}\left[ | \hat s_{t+1} - s_{t+1} |^2 + \text{MSE}(\hat r_{t+1}, r_{t+1}) \right] $$

2.2 Three key technical choices ​

ChoiceWhat it meansRepresentative
State spacePredict in raw pixels vs. predict in a low-dimensional latent spaceThe latter (Dreamer) is orders of magnitude more sample-efficient
Deterministic vs. stochasticA pure function vs. a distributional output (Gaussian / categorical)Stochastic models are more robust (don't fake certainty where there is none)
Step lengthOne-step model vs. multi-step modelOne-step models compose, but error accumulates with the number of steps

The trap in learning models

Model prediction error accumulates quadratically with planning depth: an ε error at step 1 can be amplified to the order of $n^2\varepsilon$ by step n. So "is the model accurate?" must be judged by its multi-step rollout ability, not its one-step loss — this is standard practice when evaluating world models.

3. MPC: Receding-Horizon Planning ​

3.1 The idea: skip learning a full policy — plan on the spot at every step ​

Model Predictive Control (MPC): at state $s_t$, use the model to roll forward H steps in your head, search or optimize for an action sequence, execute just the first action, then re-plan at the next moment (receding).

text
The MPC loop (runs at every timestep):
┌───────────────────────────────────────────────────────┐
│ 1. Current state s_t                                  │
│ 2. Unroll H-step candidate action sequences           │
│    inside the model                                   │
│ 3. Pick the sequence with the highest cumulative      │
│    return, a_t..a_{t+H}                               │
│ 4. Execute only the first action a_t                  │
│ 5. Observe the real s_{t+1}, then go back to 1        │
│    (the window recedes)                               │
└───────────────────────────────────────────────────────┘

3.2 What algorithm does the planning? ​

PlannerMechanismBest suited for
Random shootingRandomly sample a large batch of action sequences and take the bestLow-dimensional, simple problems
CEM (cross-entropy method)Iteratively refine the action distributionModerate complexity
MCTSTree search (see AlphaGo and Monte Carlo Tree Search)Discrete actions, sparse rewards
Gradient optimizationTake gradients through the action sequenceDifferentiable models

3.3 What MPC buys you — and what it costs ​

  • Pros: no policy network to train; adapts on the fly (re-planning every step can locally correct model errors); handles constraints natively.
  • Cons: expensive planning at every single step; a finite H means myopia (short planning horizon); still sensitive to model error.

MPC is the canonical way to trade computation for samples — it burns inference compute, not training data.

4. Dyna: Interleaving Planning and Learning ​

4.1 The idea: alternate real interaction with "mental rehearsal" ​

Dyna (Sutton, 1991) puts learning and planning into the same loop:

text
           ┌──────────────────────────────────────────────┐
           │               Dyna framework                 │
           │                                              │
  real env ──▶ direct RL learning (Q-learning)            │
     │                                                    │
     ▼                                                    │
  update world model ◀────────────────────── experience   │
     │                                                    │
     └──▶ generate simulated experience                   │
          ──▶ learn again                                 │
          (mental rehearsal × n)                          │
           └──────────────────────────────────────────────┘

Pseudocode:

text
loop:
  s, a, r, s' = interact with the environment for one step
  Q  ← Q + α(r + γ max Q(s') - Q)      # direct learning (TD)
  update model P(s'|s,a), R(s,a)
  repeat n times:                       # mental rehearsal
    (s̃, ã) ← sample a state-action pair from the model
    Q ← Q + α(r̂ + γ max Q(s̃') - Q)    # learn from simulated experience

4.2 The intuition behind Dyna, plus its variants ​

The intuition: real experience is expensive; simulated experience is almost free. Every piece of real experience triggers a model update, and then the model "reviews" it many times over. Dyna-Q is the tabular version; modern variants include Dyna-2 and neural Dyna-style methods (their thinking is continuous with Dreamer).

The philosophy of Dyna

Dyna frames "learning, planning, and acting as a trinity": learn the values (Q), learn the model (P), and plan with the model — the three reinforce one another within a single loop. It is the intellectual root of every later "model-assisted" method.

5. Latent-Space World Models: World Models, Dreamer, TD-MPC ​

5.1 World Models (Ha & Schmidhuber, 2018): compress first, then predict ​

The classic World Models paper is a three-piece design:

text
World Models: the three pieces
┌──────────────┐   ┌──────────────┐   ┌──────────────┐
│  Vision (V)  │   │  Memory (M)  │   │  Controller  │
│  autoencoder │   │ RNN predictor│   │  (policy)    │
│  pixels ──▶  │   │ predicts in  │   │ trained in   │
│  latent z    │   │ latent space │   │ the "dream"  │
└──────────────┘   └──────────────┘   └──────────────┘
  • V: compresses high-dimensional pixels into a low-dimensional latent z;
  • M: an RNN that predicts forward in latent space (this is the world model);
  • C: a policy trained inside the "dreams" the model generates — sometimes with no real environment at all.

The contribution: replacing "predict in pixels" with "predict in latent space" slashes the cost of training the policy — the policy can grind through millions of steps inside a hallucinated environment.

5.2 Dreamer: training Actor-Critic inside the learned model's "dreams" ​

Dreamer (Hafner et al., 2020/2021) fully wires a latent-space world model together with Actor-Critic:

text
The Dreamer training loop
┌───────────────────────────────────────────────────────┐
│ 1. Collect experience (real environment + current     │
│    policy)                                            │
│ 2. Learn the world model (latent dynamics + reward    │
│    predictor)                                         │
│ 3. Imagine H-step trajectories with the world model   │
│    ("dreaming")                                       │
│ 4. Train the Actor (policy) and the Critic on the     │
│    imagined trajectories                              │
│ 5. Back to 1 — an endless "real → dream → learn"      │
│    cycle                                              │
└───────────────────────────────────────────────────────┘

The Dreamer family's data efficiency leaves model-free methods far behind: on Atari, a few million frames reach the level DQN needed tens of millions of frames for; DreamerV3 was the first single agent to obtain diamonds in Minecraft, a long-horizon sparse-reward task. The key trick: the latent-space RSSM (Recurrent State-Space Model), which combines deterministic and stochastic latent states to get both long-term memory and uncertainty modeling.

5.3 TD-MPC: a TD + MPC hybrid ​

TD-MPC (Hansen et al., 2022) is an important recent entry in the "hybrid" school: it uses TD learning (a value / action-value model) to guide MPC, while keeping MPC's receding-horizon planning:

  • Learn a latent state model + a Q function (with a TD target);
  • During planning, the learned Q serves as an "estimate of far-away value," making up for MPC's limited horizon H;
  • The payoff: MPC handles precise short-term planning, the TD value backs up the long term → sample efficiency and performance in one package.

TD-MPC2 (Hansen et al., 2023) scaled this idea to SOTA across 80+ continuous-control tasks, and it is the de facto frontier of model-based continuous control today. See Frontier Progress for details.

6. The "Model-Error Trap" (Imagination Bias): Why Model-Based RL Blows Up ​

6.1 The essence of the problem: planning amplifies the model's mistakes ​

text
Real env:  s ──▶ a ──▶ s'   (true)
Model:     s ──▶ a ──▶ ŝ'   (erroneous)
                      │
            the agent keeps planning from ŝ'
                      │
            the "optimal action" is built on a wrong future
                      ▼
  behavior compounds errors in new territory (imagination bias)

This bites hardest when the data distribution the agent explores drifts away from the distribution the model was trained on — the newly learned policy visits places the model has never seen. The model starts "talking nonsense," and the policy planned on top of it collapses in the real environment.

6.2 Mitigations ​

MitigationMechanism
Plan only in trusted regionsFall back to model-free whenever model uncertainty gets too high (e.g., M2AC)
Model ensemblesTrain multiple models and downweight where they disagree most
Model uncertainty estimationUse variance / disagreement as a confidence score
Mix short and long horizonsShort-sighted MPC + long-horizon TD (the TD-MPC idea)
Distill the model into a policyLearn a model → train a policy inside it → distill it into a pure model-free policy for deployment

The most classic pitfall

"The model is accurate on the training distribution" ≠ "the model is accurate along exploration paths." Evaluate the model on the real trajectory distribution it will see at deployment, not on the training set. This is the most easily overlooked validation mistake in model-based projects.

6.3 Uncertainty quantification: making the model "know what it doesn't know" ​

Mitigating imagination bias starts with knowing where the model can't be trusted. Three engineering implementations:

MethodMechanismCharacteristics
EnsembleTrain N models with different initializations; use output disagreement as the uncertaintyThe most practical option; reliably effective; N× compute cost
Stochastic-model varianceThe model outputs a distribution (Gaussian head); use the output variance directlyCheap, but underestimates out-of-distribution (OOD) uncertainty
Density / distance estimationMeasure how far the current state is from the training data (e.g., kNN distance)Intuitive; requires a sensible feature space

The standard recipe at planning time: downgrade as soon as uncertainty crosses a threshold — switch back to short-horizon MPC or a model-free fallback (the M2AC idea), and let the model drive long-horizon planning only in trusted regions. This "uncertainty gating" is a hard prerequisite for shipping model-based RL to production.

The connection to exploration

Intrinsic rewards (RND/ICM use prediction error as a novelty signal) are essentially one flavor of model uncertainty — see Exploration and Exploitation. The very same "my prediction is off" signal is consumed as a reward during exploration and as a confidence score during planning. Once you see that, model-based RL and exploration fold into one picture.

7.1 Go is, in effect, "cheat-level" model-based RL ​

In the world of AlphaGo/AlphaZero, the transition model is perfectly known — the rules of Go are the model, and every move leads to a fully determined position. So what needs learning are two other things:

  • Policy network: compresses the search — tells MCTS "which moves are worth expanding" (shrinks the search breadth);
  • Value network: compresses evaluation — tells MCTS "roughly how good this position is" (replacing deep rollouts).
text
AlphaZero's "learning × search"
┌───────────────────────────────────────────────────────┐
│ MCTS (search — exact rollouts using the rules)        │
│   Selection: pick nodes by the UCB formula (with the  │
│     policy-network prior)                             │
│   Expansion: play a stone                             │
│   Evaluation: the value network scores the position   │
│     (replacing Monte Carlo random playouts)           │
│   Backprop: propagate the value back up the tree      │
│ Training: self-play data trains the policy network    │
│   and the value network                               │
└───────────────────────────────────────────────────────┘

The key insight: search supplies exact near-term rollouts, while networks supply generalizable far-away evaluation — the two complement each other. The policy network compresses an astronomical search space down to something tractable, and the value network frees the search from having to play the game out to the end.

7.2 MuZero: learn the "rules" too ​

AlphaZero still needed the rules — the model. MuZero (Schrittwieser et al., 2020) goes one step further: even the rules are learned by neural networks — in latent space, it learns "the next hidden state, the reward, and the value," without ever touching the real board representation. MuZero therefore doesn't need to know the rules of the game at all: it plans from raw observations + actions, and works across the board on Go, Atari, chess, and shogi.

What this means for engineers

Your task may come with "imperfect rules" or even "unknown rules," but you almost always have a better model source than MuZero did: physics engines, business simulators, human expertise. As long as model error stays under control, the sample-efficiency advantage of the model-based route is overwhelming. See RL vs. Neighboring Fields for how this connects to optimal control.

8. The Frontier, and How to Choose ​

8.1 The current landscape (2020s) ​

RouteRepresentativeOne-liner
Latent-space Dreamer lineDreamerV3Learns world models from pixels; strong at both games and control
TD+MPC lineTD-MPC2Continuous-control SOTA; sample efficiency and performance at once
Learned-search lineMuZero, AlphaZeroPlans even when the rules are unknown
LLM world modelsGenie and the world-model lineLearn the world from video/text; not yet used for serious RL decision-making

8.2 Choosing between model-based and model-free ​

text
Decision tree
├── You can get an environment simulator / physics model,
│   and its error is controllable
│   └── ▶ model-based (a crushing sample-efficiency win)
├── The environment is a black box that's hard to simulate
│   (real users, real markets)
│   └── ▶ model-free (PPO/SAC — see the actor-critic page)
├── The task has explicit rules you can simulate exactly
│   (board games, scheduling)
│   └── ▶ MCTS + learning (the AlphaGo route)
└── Budget is ample and you just want a stable final policy
    └── ▶ model-based to learn the model first ──▶ distill
        into a model-free policy

An honest note on the state of industry

Model-free methods (PPO/SAC) still dominate large-scale industrial use — because most real-world scenarios offer no reliable model. The value of model-based RL is precisely its sample efficiency: when the environment is an expensive simulator (robotics, industrial processes, scientific research), it turns from "impractical" into "the only practical option." For related cases, see RL in Science and Biomedicine.

Further Reading ​

References ​

  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Chapter 8, "Planning and Learning with Tabular Methods" (Dyna).
  • Sutton, R. S. (1991). Dyna, an Integrated Architecture for Learning, Planning, and Reacting. ACM SIGART Bulletin, 2(4), 160-163.
  • Ha, D., & Schmidhuber, J. (2018). World Models. arXiv:1803.10122
  • Hafner, D., Lillicrap, T., Fischer, I., et al. (2019). Learning Latent Dynamics for Planning from Pixels. ICML. arXiv:1811.04551 (PlaNet)
  • Hafner, D., Lillicrap, T., Ba, J., & Norouzi, M. (2020). Dream to Control: Learning Behaviors by Latent Imagination. ICLR. arXiv:1912.01603 (Dreamer)
  • Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering Diverse Domains through World Models. arXiv:2301.04104 (DreamerV3)
  • Hansen, N., Wang, X., & Su, H. (2022). Temporal Difference Learning for Model Predictive Control. ICML. arXiv:2203.04955 (TD-MPC)
  • Schrittwieser, J., Antonoglou, I., Hubert, T., et al. (2020). Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. Nature, 588, 604-609. arXiv:1911.08265 (MuZero)