Appearance
Glossary
One-line pitch: this page decodes all of RL's jargon. When a paper, an interview question, or a hallway conversation throws an acronym at you, look it up here — every entry gets a one-line definition and a pointer to the full page on this site that teaches it properly.
1. How to Use This Page
Reinforcement learning has one of the densest vocabularies in machine learning: a single concept can carry a formal name, an abbreviation, and an alias or two — temporal-difference learning goes by TD, "bootstrapping", and a couple of other names depending on the paper. This glossary exists for quick lookups, not as a substitute for the main text.
- How to read it: browse by category, or hit Ctrl+F for an abbreviation. Each entry follows the format "term | one-line definition | related page", where the related page is this site's full treatment of the topic.
- Using the related pages: follow the links when an interview follow-up or a subtle distinction calls for depth — the full pages add formulas, intuition, examples, and the usual pitfalls.
- How the concepts fit together: we recommend reading in the order of the learning path. This glossary is closer to a dictionary — look things up when stuck, then return to the main line.
- Math definitions: this page gives text-only definitions. For the formulas behind them (Bellman equation, KL divergence, expectations, and the like), see the Math Primer.
A suggestion
Treat this page as a mock-interview checklist: pick 10 entries at random — can you state each definition without looking? Whatever you can't recall is your weak spot; click through to the related page and patch it.
2. Foundational Terms
This group is the language of MDPs themselves — almost every RL problem draws on it.
| Term | One-line definition | Related page |
|---|---|---|
| Markov Decision Process (MDP) | A unified framework for sequential decision-making, written as the five-tuple (S, A, P, R, γ); nearly every RL problem can be cast as an MDP. | MDP |
| Markov Property | The property that the next state depends only on the current state and action, not on earlier history — the precondition for MDPs and the Bellman equation. | MDP |
| State | What the agent observes about the environment, s ∈ S — the basis for every decision. Whether the state representation is "Markov enough" largely determines how hard learning is. | MDP |
| Action | The choices available to the agent in each state, a ∈ A. Whether actions are discrete (left/right) or continuous (torque values) drives algorithm selection. | MDP |
| Reward | Immediate feedback r from the environment in response to an action; maximizing cumulative reward is the entire learning objective of RL. | What Is Reinforcement Learning |
| Return | Discounted cumulative reward from time t: G_t = r_t + γr_{t+1} + γ²r_{t+2} + ⋯ — the quantity value functions estimate. | MDP |
| Discount Factor (γ) | A constant 0 ≤ γ ≤ 1 that weights future rewards. Setting γ < 1 also keeps returns bounded over infinite horizons and makes convergence provable. | MDP |
| Policy | The mapping from states to action (distributions), π(a|s). It is the final product of RL training and is used directly to make decisions at deployment. | MDP |
| Stochastic vs. Deterministic Policy | A stochastic policy outputs a distribution over actions (exploration built in); a deterministic policy outputs a single action (common in continuous control). | Policy Gradient |
| State-Value Function V(s) | The expected return from state s when acting under policy π, V_π(s) — it answers "how good is this situation?" | Value Learning |
| Action-Value Function Q(s,a) | The expected return from taking action a in state s and then following policy π, Q_π(s,a) — this is what DQN learns. | Value Learning |
| Advantage Function A(s,a) | A = Q(s,a) − V(s): how much better this action is than average — the core tool for variance reduction in policy gradients. | Policy Gradient |
| Bellman Equation | The self-consistency equation for value functions, V(s) = r + γV(s′) — it decomposes long-term return into "immediate reward + next-step value". | MDP |
| Bellman Optimality | The optimal value satisfies V*(s) = max_a [r + γV*(s′)] — the mathematical basis for value iteration and Q-learning. | Value Learning |
| Exploration–Exploitation | The trade-off between trying new actions (exploration) and using the best known action (exploitation) — the first-order tension of RL. | Exploration & Exploitation |
| Regret | The gap between cumulative reward and the ideal reward from always picking the best action — the standard metric for evaluating bandit algorithms and online learning. | Multi-Armed Bandits |
| Multi-Armed Bandit | A stripped-down RL problem with one-step decisions and no time dimension — the ideal testbed for exploration–exploitation research. | Multi-Armed Bandits |
| Contextual Bandit | A bandit with context features: observe the features first, then act — the most common "RL precursor" in recommender and ad systems. | Multi-Armed Bandits |
| Credit Assignment | The problem of tracing delayed rewards back to the early actions that caused them — the core reason RL is harder than single-step supervised learning. | What Is Reinforcement Learning |
| Sparse Reward | A setting where rewards are zero for most steps and arrive only at key events (e.g., reaching the goal), demanding countermeasures like curriculum learning and intrinsic rewards. | Reward Engineering |
| Reward Shaping | Speeding up learning by adding guiding rewards (e.g., bonus for moving closer to the goal); must follow the potential-based shaping theorem, or it can change the optimal policy. | Reward Engineering |
| Reward Hacking | The agent finds and exploits loopholes in the reward function to score points without actually completing the task — e.g., gaming the metric instead of reaching the goal. | Reward Engineering |
| POMDP (Partially Observable MDP) | The agent sees only a projection of the state, not the full state (e.g., a robot with a partial map), so memory or recurrent networks are usually needed. | MDP |
INFO
It's easy to overlook the gap between POMDPs and MDPs: in engineering you almost always work with "state observations", while algorithm theory assumes "states are known". When observations are incomplete, the crudest fix is to concatenate history into the state (e.g., frame stacking).
3. Algorithm Family Terms
Arranged along the lineage "value learning → policy learning → convergence of the two" so you can compare them side by side.
| Term | One-line definition | Related page |
|---|---|---|
| Dynamic Programming (DP) | A family of methods that solve MDPs by iterating "policy evaluation + policy improvement" — applicable only when the environment model is known. | Value Learning |
| Policy Iteration | One DP method: repeatedly evaluate the current policy's value, then greedily improve it, until the policy stops changing. | Value Learning |
| Value Iteration | A DP method that iterates the Bellman optimality equation directly to converge to V*, then extracts a greedy policy from V*. | Value Learning |
| Monte Carlo (MC) | Estimates value from sample returns of complete episodes — unbiased but high-variance, and updates only happen at episode end. | Value Learning |
| Temporal-Difference (TD) Learning | Updates from "immediate reward + estimate of the next-step value" — bootstrapping. Biased but low-variance, and it updates step by step. | Value Learning |
| TD Error | δ = r + γV(s′) − V(s), the update signal for TD learning; GAE is essentially a weighted combination of multi-step TD errors. | Value Learning |
| Q-learning | The classic off-policy tabular algorithm: updates Q with max_{a'} Q(s′,a′), so the target and behavior policies can differ. | Value Learning |
| SARSA | An on-policy tabular TD algorithm that updates using the action actually taken next — it learns Q for the behavior policy itself. | Value Learning |
| On-policy / Off-policy | Whether training data comes from the same policy as the target: on-policy is more stable but sample-hungry; off-policy can reuse historical data (replay). | Value Learning |
| Experience Replay | Storing past (s, a, r, s′) transitions in a buffer and sampling them randomly to break sample correlation — one of DQN's two pillars. | Value Learning |
| Target Network | A lagging copy of the Q network that generates stable bootstrapping targets and curbs divergence — DQN's other pillar. | Value Learning |
| DQN | Neural Q-approximation + replay + target network. Beat human performance on 49 Atari games in 2015 and opened the deep RL era. | Value Learning |
| Double DQN | Decouples "choosing the action" from "evaluating the action" with two networks, relieving the systematic Q overestimation. | Value Learning |
| Dueling DQN | Splits Q into two output heads — state value V and advantage A — making value estimates more stable and learning faster. | Value Learning |
| Prioritized Replay | Weights replay sampling by TD-error magnitude, so the samples that are "not yet learned" get replayed more often. | Value Learning |
| Rainbow | Combines seven improvements — Double, Dueling, prioritized replay, multi-step returns, distributional Q (C51), noisy networks, and more — into a single agent. | Value Learning |
| Policy Gradient | Differentiates the policy parameters directly: raise the probability of high-return actions and push down low-return ones — naturally supports stochastic policies. | Policy Gradient |
| REINFORCE | The plainest policy-gradient algorithm: uses the full-episode return as the weight — clear intuition, but enormous variance. | Policy Gradient |
| Baseline | A quantity subtracted from the return (usually the value function V); it leaves the gradient expectation unchanged while slashing variance. | Policy Gradient |
| Actor-Critic | A pairing of a policy network (actor) and a value network (critic); the critic supplies the actor with a low-variance advantage signal — the dominant architecture today. | Actor-Critic Family |
| A2C / A3C | Actor-Critic running across parallel environments: A3C is the asynchronous version, A2C the synchronous one (more common and easier to reproduce). | Actor-Critic Family |
| GAE (Generalized Advantage Estimation) | Weights n-step advantages with a parameter λ, trading bias against variance along a continuous spectrum — PPO's standard companion. | Actor-Critic Family |
| TRPO | Caps each policy update with a KL-divergence constraint, guaranteeing monotonic improvement — but complex to implement and computationally heavy. | Policy Gradient |
| PPO | Approximates TRPO's constraint with a clipped objective — simple, stable, easy to implement, and the most mainstream on-policy algorithm today. | Policy Gradient |
| DDPG | Deterministic policy gradients + Q-learning: an off-policy continuous-control algorithm and the early form of "deep actor-critic". | Actor-Critic Family |
| TD3 | Three fixes to DDPG: twin Q networks, delayed policy updates, and target policy smoothing — greatly relieving value overestimation and variance. | Actor-Critic Family |
| SAC | A maximum-entropy off-policy algorithm whose objective additionally maximizes policy entropy — one of the de facto standards for continuous control. | Actor-Critic Family |
| Entropy Regularization | Adds a policy-entropy term to the objective to encourage randomness, preventing premature convergence and preserving exploration. | Exploration & Exploitation |
| Temperature Coefficient (α) | In SAC, the coefficient weighting the entropy bonus; it can be tuned automatically (auto α adjustment), sparing you the manual search. | Actor-Critic Family |
A classic gotcha
SARSA vs. Q-learning: the two formulas differ by exactly one choice — SARSA uses the a′ actually taken, Q-learning uses the a′ picked by max. That single choice determines the on/off-policy property, and it comes up again and again in interviews and debugging. See Value Learning.
4. Deep RL Engineering Terms
High-frequency vocabulary for the "make it actually work" phase — mostly about why DQN/PPO work and how to deploy them.
| Term | One-line definition | Related page |
|---|---|---|
| Sample Efficiency | How many environment interactions it takes to reach a given performance; off-policy algorithms (SAC, DQN) are typically far more sample-efficient than on-policy ones (PPO). | Actor-Critic Family |
| Learning Curve | A plot of return against training steps / environment interactions — the first diagnostic tool for RL tuning and evaluation. | Evaluation & Benchmarks |
| Seed | The random seed; fixing it makes experiments reproducible, but tuning repeatedly against the same seed causes hidden overfitting. | Evaluation & Benchmarks |
| Reproducibility | Whether the same code and config can reproduce similar results — requires fixed seeds, fully logged hyperparameters, and pinned environment versions. | Common Pitfalls & Anti-Patterns |
| Distribution Shift | Mismatch between training and deployment distributions (e.g., sim → real, offline → online); RL policies often collapse because of it. | Offline RL |
| Domain Randomization | Randomizing simulated physics/rendering parameters during training so the policy learns invariances — a key enabler of zero-shot Sim2Real. | Robotics Sim2Real |
| Sim2Real | The transfer from training in simulation to deploying in the real world; the core gap is modeling error in dynamics, observations, and contact. | Robotics Sim2Real |
| Curriculum Learning | A training schedule that gradually moves from easy tasks to hard ones, countering sparse rewards and exploration difficulty. | Reward Engineering |
| HER (Hindsight Experience Replay) | Rewrites "goals not achieved" as "achieved" to fabricate positive samples, solving sparse-reward and multi-goal problems. | Reward Engineering |
| Intrinsic Reward / Curiosity | Uses intrinsic signals such as novelty or prediction error as extra rewards, driving the agent to explore unseen states. | Exploration & Exploitation |
| Count-Based Exploration | Counts state visits and bonuses rare states — a naive method that works in tabular worlds and fails in high dimensions. | Exploration & Exploitation |
| RND (Random Network Distillation) | Uses a fixed random network's output as a "novelty target": states with large prediction error are more "novel" and earn higher intrinsic reward. | Exploration & Exploitation |
| Vectorized Environments | Running multiple environment copies in parallel to collect experience — boosts throughput and training stability; standard engineering for A2C/PPO. | Choosing Frameworks & Tools |
5. Advanced Paradigm Terms
Vocabulary for frontier paradigms — offline RL, multi-agent, alignment — essential reading for 2020s papers.
| Term | One-line definition | Related page |
|---|---|---|
| Model-Based RL | Learns an environment model first, then plans and learns inside it — sample-efficient, but model error compounds. | Model-Based RL |
| World Model | A model that encodes state representations and predicts transitions and rewards — the core component of Dreamer, TD-MPC, and MuZero. | Model-Based RL |
| Model-Free RL | Algorithms that skip the environment model and learn value/policy directly from interaction data — DQN, PPO, and SAC are all model-free. | Value Learning |
| Offline RL | Trains only on a fixed historical dataset, with no further environment interaction; the key difficulty is value overestimation on OOD actions. | Offline RL |
| OOD (Out-of-Distribution) | State-action pairs absent from, or barely covered by, the training data; in offline RL the "hallucinated overestimates" that Q networks extrapolate for them are the root disease. | Offline RL |
| CQL (Conservative Q-Learning) | Adds a penalty on the values of OOD actions to the Q-learning objective — an offline RL method that fights overestimation with "conservatism". | Offline RL |
| IQL (Implicit Q-Learning) | Regresses values only for in-distribution actions without extrapolation — an offline RL method that needs no explicit behavior regularizer. | Offline RL |
| Behavior Cloning (BC) | Direct supervised learning from expert demonstrations (state in, action out) — simple, but unable to handle states the expert never visited. | RL vs. Neighboring Fields |
| Imitation Learning | The family of paradigms for learning policies from demonstrations: behavior cloning, inverse RL, GAIL, and more. | RL vs. Neighboring Fields |
| Inverse RL (IRL) | Infers the reward function from expert behavior — another path forward when reward engineering is too hard to write by hand. | Reward Engineering |
| Multi-Agent RL (MARL) | Multiple agents learn to make decisions simultaneously; the environment dynamics shift as others' policies change, making the problem non-stationary. | Multi-Agent RL |
| CTDE | "Centralized training, decentralized execution": global information is available during training, but deployment uses only local observations — the mainstream MARL paradigm. | Multi-Agent RL |
| Nash Equilibrium | A strategy profile where no agent can improve its payoff by unilaterally deviating — a convergence target for MARL that is often also where the difficulty lies. | Multi-Agent RL |
| Non-Stationarity | The environment dynamics keep changing because other agents are learning at the same time, so the single-agent MDP assumption breaks down. | Multi-Agent RL |
| Value Decomposition | Decomposes the joint Q value into a sum or product of individual Q values (e.g., QMIX) — the classic way to realize CTDE. | Multi-Agent RL |
| MADDPG | A multi-agent extension of DDPG with centralized critics and decentralized actors; off-policy, continuous actions. | Multi-Agent RL |
| MAPPO | The multi-agent version of PPO; simple and direct, yet it has set new SOTA results on many tasks. | Multi-Agent RL |
| RLHF | A three-stage alignment pipeline: train a reward model from human preferences, then fine-tune the language model with PPO (SFT → RM → PPO). | RLHF & Human-Feedback Alignment |
| Reward Model | A model trained on pairs of human preferences to score candidate text — it serves as "automated human feedback" in RLHF. | RLHF & Human-Feedback Alignment |
| Bradley–Terry Model | The standard preference-modeling assumption: P(choose A over B) = σ(r(A) − r(B)) — the foundation of reward-model training. | RLHF & Human-Feedback Alignment |
| KL Penalty | A regularizer in RLHF that keeps the new policy from drifting too far from the reference model, preventing PPO from "flying off". | RLHF & Human-Feedback Alignment |
| DPO (Direct Preference Optimization) | An alignment shortcut that builds the policy objective directly from preference data, skipping both the reward model and PPO. | RLHF & Human-Feedback Alignment |
| Alignment | The overarching goal of making model behavior match human intent and values; RLHF, DPO, and RLAIF are all means to it. | RLHF & Human-Feedback Alignment |
| Alignment Tax | The price of alignment training, paid in degraded general capability — the classic side effect of RLHF. | RLHF & Human-Feedback Alignment |
| Reward Over-Optimization | The phenomenon where the reward-model score climbs while true quality drops (Goodhart's law) — a typical failure mode of RLHF training. | RLHF & Human-Feedback Alignment |
| RLAIF | Replaces human feedback with AI feedback to train the reward model, scaling alignment up. | RLHF & Human-Feedback Alignment |
6. Easily Confused Terms
The five-plus confusion clusters below are the most common "danger zones" in interviews and paper reading, pulled out here for side-by-side comparison.
| Confusable pair | The key distinction | One-line memory hook |
|---|---|---|
| MC vs. TD | MC uses full-episode returns (unbiased, high variance); TD uses "reward + bootstrapped estimate" (biased, low variance) | "MC waits for the outcome; TD guesses one step" |
| On-policy vs. Off-policy | Whether the data comes from the target policy; off-policy can reuse historical data | "SARSA learns from itself; Q-learning learns from others" |
| Value Learning vs. Policy Gradient | Value learning learns Q first, then derives a policy; policy gradient learns the policy directly | "Estimate-then-choose vs. choose directly" |
| Model-Based vs. Model-Free | Whether an explicit model of the transition/reward dynamics is learned | "Do you need to learn the world first?" |
| RLHF vs. DPO | RLHF has three stages (RM + PPO); DPO is one-step, with no RM and no PPO | "The detour vs. the direct route" |
| TRPO vs. PPO | TRPO enforces a hard KL constraint; PPO replaces it with a soft clip approximation | "A constraint turned into a penalty (approximation)" |
| DDPG vs. TD3 vs. SAC | All three are off-policy continuous control; TD3 fixes DDPG's overestimation, and SAC adds maximum entropy on top | "SAC = TD3's ideas + entropy" |
7. Abbreviations A–Z
Every abbreviation on this page, alphabetized — when you can't remember where one came from, look it up here.
| Abbrev. | Full name | Where on this page |
|---|---|---|
| A2C / A3C | Advantage Actor-Critic / Asynchronous A2C | Algorithm Families |
| AC | Actor-Critic | Algorithm Families |
| BC | Behavior Cloning | Advanced Paradigms |
| BT | Bradley–Terry preference model | Advanced Paradigms |
| CQL | Conservative Q-Learning | Advanced Paradigms |
| CTDE | Centralized Training with Decentralized Execution | Advanced Paradigms |
| DDPG | Deep Deterministic Policy Gradient | Algorithm Families |
| DPO | Direct Preference Optimization | Advanced Paradigms |
| DP | Dynamic Programming | Algorithm Families |
| DQN | Deep Q-Network | Algorithm Families |
| GAE | Generalized Advantage Estimation | Algorithm Families |
| HER | Hindsight Experience Replay | Engineering Terms |
| IQL | Implicit Q-Learning | Advanced Paradigms |
| IRL | Inverse Reinforcement Learning | Advanced Paradigms |
| KL | Kullback–Leibler divergence | Math Primer |
| MARL | Multi-Agent RL | Advanced Paradigms |
| MC | Monte Carlo | Algorithm Families |
| MDP | Markov Decision Process | Foundations |
| OOD | Out-of-Distribution | Advanced Paradigms |
| POMDP | Partially Observable MDP | Foundations |
| PPO | Proximal Policy Optimization | Algorithm Families |
| RM | Reward Model | Advanced Paradigms |
| RND | Random Network Distillation | Engineering Terms |
| RLHF | RL from Human Feedback | Advanced Paradigms |
| SAC | Soft Actor-Critic | Algorithm Families |
| SARSA | State-Action-Reward-State-Action | Algorithm Families |
| SFT | Supervised Fine-Tuning | RLHF |
| TD | Temporal Difference | Algorithm Families |
| TD3 | Twin Delayed DDPG | Algorithm Families |
| TRPO | Trust Region Policy Optimization | Algorithm Families |
DANGER
Abbreviations with the same spelling but different meanings are common in papers: TD can mean "temporal difference" in one paper and "task description" in another; in some contexts IML stands for "implicit imitation learning". Always pin down an abbreviation from its context, and in interviews, state the full name before elaborating.
Further Reading
- Markov Decision Process (MDP) — the theoretical foundation of this glossary: full derivation of the five-tuple and the Bellman equation.
- Value Learning: From Dynamic Programming to DQN — a systematic treatment of the DP/MC/TD/Q-learning/DQN family.
- Policy Gradient Methods — intuition and formulas for REINFORCE, advantage, and PPO.
- RLHF & Human-Feedback Alignment — the full story of alignment terminology (RM, KL, DPO, alignment tax).
- Math Primer — the math tools behind these entries and the algorithms that use them.
- Curated Resource List — want a structured course? Pick your textbooks and courses here.
References
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed., MIT Press. Free online version: http://incompleteideas.net/book/RLbook2020.pdf — the authoritative source for these terms.
- OpenAI Spinning Up (glossary section): https://spinningup.openai.com/en/latest/ — an English-language RL terminology reference that complements this page.
- David Silver's UCL reinforcement learning course: https://www.davidsilver.uk/teaching/ — full lecture videos and slides covering these terms.
- Lilian Weng's blog: https://lilianweng.github.io/ — long-form essays on frontier terms including RLHF, offline RL, and world models.