Appearance
Model-Based RL and World Models
In one line: this page explains model-based reinforcement learning (model-based RL) — instead of learning a policy directly, you first learn the "dynamics of the environment" (a world model) and then plan your actions inside that mental simulation. After reading it, you'll be able to state the essential difference between model-free and model-based, trace the three generations of techniques from MPC through Dyna to Dreamer, recognize the "model-error trap," and reread "learning × search" through AlphaGo's eyes.
1. Model-Free vs. Model-Based: Two Fundamentally Different Routes
1.1 The core difference: whether you "understand the world"
- Model-free: never learns the world's regularities; it improves the value function or policy directly from trial-and-error samples. It answers only "what should I do?" (which action to pick), never "how does the world work?"
- Model-based: first learns "if I take action a, what will the world look like and what reward will I get?" (transition + reward), and then plans inside the learned world model. It answers one extra question: "what happens if I do this?"
text
Model-free: samples ──▶ value/policy (skips "understanding")
only learns "what to do"
Model-based: samples ──▶ world model (ŝ, r̂) ──▶ mental planning ──▶ action
first learn "how the world works", then roll it forward1.2 Side by side
| Dimension | Model-free | Model-based |
|---|---|---|
| What it learns | Q(s,a) / π(a | s) |
| Sample efficiency | Low (millions of steps and up) | High (tens of thousands of steps or fewer) |
| Where the final policy comes from | Direct optimization | Planned from, or distilled out of, the model |
| Generalization to new tasks | Poor (learn it all again) | Good (the model is reusable and re-plannable) |
| Main risk | Burning through samples | Model error (imagination bias) |
| Representatives | DQN, PPO, SAC | Dyna, Dreamer, TD-MPC, MuZero |
The one-line intuition
Model-free is like "grinding practice problems until test-taking instinct kicks in"; model-based is like "learning the physics first and then solving the problems." The former burns a lot of samples but is blunt and simple; the latter saves samples, but "if you learned the physics wrong, everything downstream is wrong."
2. Learning a World Model: Transition + Reward
2.1 Formalization: the forward model
A world model predicts the outcome of a given state and action:
$$ \hat s_{t+1} = f_\phi(s_t, a_t), \qquad \hat r_{t+1} = g_\psi(s_t, a_t) $$
The training objective is plain supervised learning: on real interaction data $(s_t, a_t, s_{t+1}, r_{t+1})$, minimize the prediction error.
$$ \mathcal{L} = \mathbb{E}\left[ | \hat s_{t+1} - s_{t+1} |^2 + \text{MSE}(\hat r_{t+1}, r_{t+1}) \right] $$
2.2 Three key technical choices
| Choice | What it means | Representative |
|---|---|---|
| State space | Predict in raw pixels vs. predict in a low-dimensional latent space | The latter (Dreamer) is orders of magnitude more sample-efficient |
| Deterministic vs. stochastic | A pure function vs. a distributional output (Gaussian / categorical) | Stochastic models are more robust (don't fake certainty where there is none) |
| Step length | One-step model vs. multi-step model | One-step models compose, but error accumulates with the number of steps |
The trap in learning models
Model prediction error accumulates quadratically with planning depth: an ε error at step 1 can be amplified to the order of $n^2\varepsilon$ by step n. So "is the model accurate?" must be judged by its multi-step rollout ability, not its one-step loss — this is standard practice when evaluating world models.
3. MPC: Receding-Horizon Planning
3.1 The idea: skip learning a full policy — plan on the spot at every step
Model Predictive Control (MPC): at state $s_t$, use the model to roll forward H steps in your head, search or optimize for an action sequence, execute just the first action, then re-plan at the next moment (receding).
text
The MPC loop (runs at every timestep):
┌───────────────────────────────────────────────────────┐
│ 1. Current state s_t │
│ 2. Unroll H-step candidate action sequences │
│ inside the model │
│ 3. Pick the sequence with the highest cumulative │
│ return, a_t..a_{t+H} │
│ 4. Execute only the first action a_t │
│ 5. Observe the real s_{t+1}, then go back to 1 │
│ (the window recedes) │
└───────────────────────────────────────────────────────┘3.2 What algorithm does the planning?
| Planner | Mechanism | Best suited for |
|---|---|---|
| Random shooting | Randomly sample a large batch of action sequences and take the best | Low-dimensional, simple problems |
| CEM (cross-entropy method) | Iteratively refine the action distribution | Moderate complexity |
| MCTS | Tree search (see AlphaGo and Monte Carlo Tree Search) | Discrete actions, sparse rewards |
| Gradient optimization | Take gradients through the action sequence | Differentiable models |
3.3 What MPC buys you — and what it costs
- Pros: no policy network to train; adapts on the fly (re-planning every step can locally correct model errors); handles constraints natively.
- Cons: expensive planning at every single step; a finite H means myopia (short planning horizon); still sensitive to model error.
MPC is the canonical way to trade computation for samples — it burns inference compute, not training data.
4. Dyna: Interleaving Planning and Learning
4.1 The idea: alternate real interaction with "mental rehearsal"
Dyna (Sutton, 1991) puts learning and planning into the same loop:
text
┌──────────────────────────────────────────────┐
│ Dyna framework │
│ │
real env ──▶ direct RL learning (Q-learning) │
│ │
▼ │
update world model ◀────────────────────── experience │
│ │
└──▶ generate simulated experience │
──▶ learn again │
(mental rehearsal × n) │
└──────────────────────────────────────────────┘Pseudocode:
text
loop:
s, a, r, s' = interact with the environment for one step
Q ← Q + α(r + γ max Q(s') - Q) # direct learning (TD)
update model P(s'|s,a), R(s,a)
repeat n times: # mental rehearsal
(s̃, ã) ← sample a state-action pair from the model
Q ← Q + α(r̂ + γ max Q(s̃') - Q) # learn from simulated experience4.2 The intuition behind Dyna, plus its variants
The intuition: real experience is expensive; simulated experience is almost free. Every piece of real experience triggers a model update, and then the model "reviews" it many times over. Dyna-Q is the tabular version; modern variants include Dyna-2 and neural Dyna-style methods (their thinking is continuous with Dreamer).
The philosophy of Dyna
Dyna frames "learning, planning, and acting as a trinity": learn the values (Q), learn the model (P), and plan with the model — the three reinforce one another within a single loop. It is the intellectual root of every later "model-assisted" method.
5. Latent-Space World Models: World Models, Dreamer, TD-MPC
5.1 World Models (Ha & Schmidhuber, 2018): compress first, then predict
The classic World Models paper is a three-piece design:
text
World Models: the three pieces
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Vision (V) │ │ Memory (M) │ │ Controller │
│ autoencoder │ │ RNN predictor│ │ (policy) │
│ pixels ──▶ │ │ predicts in │ │ trained in │
│ latent z │ │ latent space │ │ the "dream" │
└──────────────┘ └──────────────┘ └──────────────┘- V: compresses high-dimensional pixels into a low-dimensional latent z;
- M: an RNN that predicts forward in latent space (this is the world model);
- C: a policy trained inside the "dreams" the model generates — sometimes with no real environment at all.
The contribution: replacing "predict in pixels" with "predict in latent space" slashes the cost of training the policy — the policy can grind through millions of steps inside a hallucinated environment.
5.2 Dreamer: training Actor-Critic inside the learned model's "dreams"
Dreamer (Hafner et al., 2020/2021) fully wires a latent-space world model together with Actor-Critic:
text
The Dreamer training loop
┌───────────────────────────────────────────────────────┐
│ 1. Collect experience (real environment + current │
│ policy) │
│ 2. Learn the world model (latent dynamics + reward │
│ predictor) │
│ 3. Imagine H-step trajectories with the world model │
│ ("dreaming") │
│ 4. Train the Actor (policy) and the Critic on the │
│ imagined trajectories │
│ 5. Back to 1 — an endless "real → dream → learn" │
│ cycle │
└───────────────────────────────────────────────────────┘The Dreamer family's data efficiency leaves model-free methods far behind: on Atari, a few million frames reach the level DQN needed tens of millions of frames for; DreamerV3 was the first single agent to obtain diamonds in Minecraft, a long-horizon sparse-reward task. The key trick: the latent-space RSSM (Recurrent State-Space Model), which combines deterministic and stochastic latent states to get both long-term memory and uncertainty modeling.
5.3 TD-MPC: a TD + MPC hybrid
TD-MPC (Hansen et al., 2022) is an important recent entry in the "hybrid" school: it uses TD learning (a value / action-value model) to guide MPC, while keeping MPC's receding-horizon planning:
- Learn a latent state model + a Q function (with a TD target);
- During planning, the learned Q serves as an "estimate of far-away value," making up for MPC's limited horizon H;
- The payoff: MPC handles precise short-term planning, the TD value backs up the long term → sample efficiency and performance in one package.
TD-MPC2 (Hansen et al., 2023) scaled this idea to SOTA across 80+ continuous-control tasks, and it is the de facto frontier of model-based continuous control today. See Frontier Progress for details.
6. The "Model-Error Trap" (Imagination Bias): Why Model-Based RL Blows Up
6.1 The essence of the problem: planning amplifies the model's mistakes
text
Real env: s ──▶ a ──▶ s' (true)
Model: s ──▶ a ──▶ ŝ' (erroneous)
│
the agent keeps planning from ŝ'
│
the "optimal action" is built on a wrong future
▼
behavior compounds errors in new territory (imagination bias)This bites hardest when the data distribution the agent explores drifts away from the distribution the model was trained on — the newly learned policy visits places the model has never seen. The model starts "talking nonsense," and the policy planned on top of it collapses in the real environment.
6.2 Mitigations
| Mitigation | Mechanism |
|---|---|
| Plan only in trusted regions | Fall back to model-free whenever model uncertainty gets too high (e.g., M2AC) |
| Model ensembles | Train multiple models and downweight where they disagree most |
| Model uncertainty estimation | Use variance / disagreement as a confidence score |
| Mix short and long horizons | Short-sighted MPC + long-horizon TD (the TD-MPC idea) |
| Distill the model into a policy | Learn a model → train a policy inside it → distill it into a pure model-free policy for deployment |
The most classic pitfall
"The model is accurate on the training distribution" ≠ "the model is accurate along exploration paths." Evaluate the model on the real trajectory distribution it will see at deployment, not on the training set. This is the most easily overlooked validation mistake in model-based projects.
6.3 Uncertainty quantification: making the model "know what it doesn't know"
Mitigating imagination bias starts with knowing where the model can't be trusted. Three engineering implementations:
| Method | Mechanism | Characteristics |
|---|---|---|
| Ensemble | Train N models with different initializations; use output disagreement as the uncertainty | The most practical option; reliably effective; N× compute cost |
| Stochastic-model variance | The model outputs a distribution (Gaussian head); use the output variance directly | Cheap, but underestimates out-of-distribution (OOD) uncertainty |
| Density / distance estimation | Measure how far the current state is from the training data (e.g., kNN distance) | Intuitive; requires a sensible feature space |
The standard recipe at planning time: downgrade as soon as uncertainty crosses a threshold — switch back to short-horizon MPC or a model-free fallback (the M2AC idea), and let the model drive long-horizon planning only in trusted regions. This "uncertainty gating" is a hard prerequisite for shipping model-based RL to production.
The connection to exploration
Intrinsic rewards (RND/ICM use prediction error as a novelty signal) are essentially one flavor of model uncertainty — see Exploration and Exploitation. The very same "my prediction is off" signal is consumed as a reward during exploration and as a confidence score during planning. Once you see that, model-based RL and exploration fold into one picture.
7. The AlphaGo Lens: A Perfect Model + Learned Search
7.1 Go is, in effect, "cheat-level" model-based RL
In the world of AlphaGo/AlphaZero, the transition model is perfectly known — the rules of Go are the model, and every move leads to a fully determined position. So what needs learning are two other things:
- Policy network: compresses the search — tells MCTS "which moves are worth expanding" (shrinks the search breadth);
- Value network: compresses evaluation — tells MCTS "roughly how good this position is" (replacing deep rollouts).
text
AlphaZero's "learning × search"
┌───────────────────────────────────────────────────────┐
│ MCTS (search — exact rollouts using the rules) │
│ Selection: pick nodes by the UCB formula (with the │
│ policy-network prior) │
│ Expansion: play a stone │
│ Evaluation: the value network scores the position │
│ (replacing Monte Carlo random playouts) │
│ Backprop: propagate the value back up the tree │
│ Training: self-play data trains the policy network │
│ and the value network │
└───────────────────────────────────────────────────────┘The key insight: search supplies exact near-term rollouts, while networks supply generalizable far-away evaluation — the two complement each other. The policy network compresses an astronomical search space down to something tractable, and the value network frees the search from having to play the game out to the end.
7.2 MuZero: learn the "rules" too
AlphaZero still needed the rules — the model. MuZero (Schrittwieser et al., 2020) goes one step further: even the rules are learned by neural networks — in latent space, it learns "the next hidden state, the reward, and the value," without ever touching the real board representation. MuZero therefore doesn't need to know the rules of the game at all: it plans from raw observations + actions, and works across the board on Go, Atari, chess, and shogi.
What this means for engineers
Your task may come with "imperfect rules" or even "unknown rules," but you almost always have a better model source than MuZero did: physics engines, business simulators, human expertise. As long as model error stays under control, the sample-efficiency advantage of the model-based route is overwhelming. See RL vs. Neighboring Fields for how this connects to optimal control.
8. The Frontier, and How to Choose
8.1 The current landscape (2020s)
| Route | Representative | One-liner |
|---|---|---|
| Latent-space Dreamer line | DreamerV3 | Learns world models from pixels; strong at both games and control |
| TD+MPC line | TD-MPC2 | Continuous-control SOTA; sample efficiency and performance at once |
| Learned-search line | MuZero, AlphaZero | Plans even when the rules are unknown |
| LLM world models | Genie and the world-model line | Learn the world from video/text; not yet used for serious RL decision-making |
8.2 Choosing between model-based and model-free
text
Decision tree
├── You can get an environment simulator / physics model,
│ and its error is controllable
│ └── ▶ model-based (a crushing sample-efficiency win)
├── The environment is a black box that's hard to simulate
│ (real users, real markets)
│ └── ▶ model-free (PPO/SAC — see the actor-critic page)
├── The task has explicit rules you can simulate exactly
│ (board games, scheduling)
│ └── ▶ MCTS + learning (the AlphaGo route)
└── Budget is ample and you just want a stable final policy
└── ▶ model-based to learn the model first ──▶ distill
into a model-free policyAn honest note on the state of industry
Model-free methods (PPO/SAC) still dominate large-scale industrial use — because most real-world scenarios offer no reliable model. The value of model-based RL is precisely its sample efficiency: when the environment is an expensive simulator (robotics, industrial processes, scientific research), it turns from "impractical" into "the only practical option." For related cases, see RL in Science and Biomedicine.
Further Reading
- Markov Decision Processes (MDP) — what model-based RL sets out to learn is exactly the P and R of an MDP
- AlphaGo and Monte Carlo Tree Search — the full case study of a perfect model plus learned search
- Frontier Progress — the latest on DreamerV3, TD-MPC2, and the MuZero successors
- RL vs. Neighboring Fields — how MPC relates to optimal control, and the intellectual neighborhood of model-based RL
- RL in Science and Biomedicine — model-based RL deployed under expensive simulators
References
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Chapter 8, "Planning and Learning with Tabular Methods" (Dyna).
- Sutton, R. S. (1991). Dyna, an Integrated Architecture for Learning, Planning, and Reacting. ACM SIGART Bulletin, 2(4), 160-163.
- Ha, D., & Schmidhuber, J. (2018). World Models. arXiv:1803.10122
- Hafner, D., Lillicrap, T., Fischer, I., et al. (2019). Learning Latent Dynamics for Planning from Pixels. ICML. arXiv:1811.04551 (PlaNet)
- Hafner, D., Lillicrap, T., Ba, J., & Norouzi, M. (2020). Dream to Control: Learning Behaviors by Latent Imagination. ICLR. arXiv:1912.01603 (Dreamer)
- Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering Diverse Domains through World Models. arXiv:2301.04104 (DreamerV3)
- Hansen, N., Wang, X., & Su, H. (2022). Temporal Difference Learning for Model Predictive Control. ICML. arXiv:2203.04955 (TD-MPC)
- Schrittwieser, J., Antonoglou, I., Hubert, T., et al. (2020). Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. Nature, 588, 604-609. arXiv:1911.08265 (MuZero)