Skip to content

Glossary

On this page Quick reference for RL jargon — MDP, Bellman equation, TD, MC, GAE, PPO, SAC, DQN, OOD, CTDE, RLHF, reward hacking, and 60+ more entries, each with a one-line definition plus a link to the page that covers it in depth.

Glossary ​

One-line pitch: this page decodes all of RL's jargon. When a paper, an interview question, or a hallway conversation throws an acronym at you, look it up here — every entry gets a one-line definition and a pointer to the full page on this site that teaches it properly.

1. How to Use This Page ​

Reinforcement learning has one of the densest vocabularies in machine learning: a single concept can carry a formal name, an abbreviation, and an alias or two — temporal-difference learning goes by TD, "bootstrapping", and a couple of other names depending on the paper. This glossary exists for quick lookups, not as a substitute for the main text.

  • How to read it: browse by category, or hit Ctrl+F for an abbreviation. Each entry follows the format "term | one-line definition | related page", where the related page is this site's full treatment of the topic.
  • Using the related pages: follow the links when an interview follow-up or a subtle distinction calls for depth — the full pages add formulas, intuition, examples, and the usual pitfalls.
  • How the concepts fit together: we recommend reading in the order of the learning path. This glossary is closer to a dictionary — look things up when stuck, then return to the main line.
  • Math definitions: this page gives text-only definitions. For the formulas behind them (Bellman equation, KL divergence, expectations, and the like), see the Math Primer.

A suggestion

Treat this page as a mock-interview checklist: pick 10 entries at random — can you state each definition without looking? Whatever you can't recall is your weak spot; click through to the related page and patch it.

2. Foundational Terms ​

This group is the language of MDPs themselves — almost every RL problem draws on it.

TermOne-line definitionRelated page
Markov Decision Process (MDP)A unified framework for sequential decision-making, written as the five-tuple (S, A, P, R, γ); nearly every RL problem can be cast as an MDP.MDP
Markov PropertyThe property that the next state depends only on the current state and action, not on earlier history — the precondition for MDPs and the Bellman equation.MDP
StateWhat the agent observes about the environment, s ∈ S — the basis for every decision. Whether the state representation is "Markov enough" largely determines how hard learning is.MDP
ActionThe choices available to the agent in each state, a ∈ A. Whether actions are discrete (left/right) or continuous (torque values) drives algorithm selection.MDP
RewardImmediate feedback r from the environment in response to an action; maximizing cumulative reward is the entire learning objective of RL.What Is Reinforcement Learning
ReturnDiscounted cumulative reward from time t: G_t = r_t + γr_{t+1} + γ²r_{t+2} + ⋯ — the quantity value functions estimate.MDP
Discount Factor (γ)A constant 0 ≤ γ ≤ 1 that weights future rewards. Setting γ < 1 also keeps returns bounded over infinite horizons and makes convergence provable.MDP
PolicyThe mapping from states to action (distributions), π(a|s). It is the final product of RL training and is used directly to make decisions at deployment.MDP
Stochastic vs. Deterministic PolicyA stochastic policy outputs a distribution over actions (exploration built in); a deterministic policy outputs a single action (common in continuous control).Policy Gradient
State-Value Function V(s)The expected return from state s when acting under policy π, V_π(s) — it answers "how good is this situation?"Value Learning
Action-Value Function Q(s,a)The expected return from taking action a in state s and then following policy π, Q_π(s,a) — this is what DQN learns.Value Learning
Advantage Function A(s,a)A = Q(s,a) − V(s): how much better this action is than average — the core tool for variance reduction in policy gradients.Policy Gradient
Bellman EquationThe self-consistency equation for value functions, V(s) = r + γV(s′) — it decomposes long-term return into "immediate reward + next-step value".MDP
Bellman OptimalityThe optimal value satisfies V*(s) = max_a [r + γV*(s′)] — the mathematical basis for value iteration and Q-learning.Value Learning
Exploration–ExploitationThe trade-off between trying new actions (exploration) and using the best known action (exploitation) — the first-order tension of RL.Exploration & Exploitation
RegretThe gap between cumulative reward and the ideal reward from always picking the best action — the standard metric for evaluating bandit algorithms and online learning.Multi-Armed Bandits
Multi-Armed BanditA stripped-down RL problem with one-step decisions and no time dimension — the ideal testbed for exploration–exploitation research.Multi-Armed Bandits
Contextual BanditA bandit with context features: observe the features first, then act — the most common "RL precursor" in recommender and ad systems.Multi-Armed Bandits
Credit AssignmentThe problem of tracing delayed rewards back to the early actions that caused them — the core reason RL is harder than single-step supervised learning.What Is Reinforcement Learning
Sparse RewardA setting where rewards are zero for most steps and arrive only at key events (e.g., reaching the goal), demanding countermeasures like curriculum learning and intrinsic rewards.Reward Engineering
Reward ShapingSpeeding up learning by adding guiding rewards (e.g., bonus for moving closer to the goal); must follow the potential-based shaping theorem, or it can change the optimal policy.Reward Engineering
Reward HackingThe agent finds and exploits loopholes in the reward function to score points without actually completing the task — e.g., gaming the metric instead of reaching the goal.Reward Engineering
POMDP (Partially Observable MDP)The agent sees only a projection of the state, not the full state (e.g., a robot with a partial map), so memory or recurrent networks are usually needed.MDP

INFO

It's easy to overlook the gap between POMDPs and MDPs: in engineering you almost always work with "state observations", while algorithm theory assumes "states are known". When observations are incomplete, the crudest fix is to concatenate history into the state (e.g., frame stacking).

3. Algorithm Family Terms ​

Arranged along the lineage "value learning → policy learning → convergence of the two" so you can compare them side by side.

TermOne-line definitionRelated page
Dynamic Programming (DP)A family of methods that solve MDPs by iterating "policy evaluation + policy improvement" — applicable only when the environment model is known.Value Learning
Policy IterationOne DP method: repeatedly evaluate the current policy's value, then greedily improve it, until the policy stops changing.Value Learning
Value IterationA DP method that iterates the Bellman optimality equation directly to converge to V*, then extracts a greedy policy from V*.Value Learning
Monte Carlo (MC)Estimates value from sample returns of complete episodes — unbiased but high-variance, and updates only happen at episode end.Value Learning
Temporal-Difference (TD) LearningUpdates from "immediate reward + estimate of the next-step value" — bootstrapping. Biased but low-variance, and it updates step by step.Value Learning
TD Errorδ = r + γV(s′) − V(s), the update signal for TD learning; GAE is essentially a weighted combination of multi-step TD errors.Value Learning
Q-learningThe classic off-policy tabular algorithm: updates Q with max_{a'} Q(s′,a′), so the target and behavior policies can differ.Value Learning
SARSAAn on-policy tabular TD algorithm that updates using the action actually taken next — it learns Q for the behavior policy itself.Value Learning
On-policy / Off-policyWhether training data comes from the same policy as the target: on-policy is more stable but sample-hungry; off-policy can reuse historical data (replay).Value Learning
Experience ReplayStoring past (s, a, r, s′) transitions in a buffer and sampling them randomly to break sample correlation — one of DQN's two pillars.Value Learning
Target NetworkA lagging copy of the Q network that generates stable bootstrapping targets and curbs divergence — DQN's other pillar.Value Learning
DQNNeural Q-approximation + replay + target network. Beat human performance on 49 Atari games in 2015 and opened the deep RL era.Value Learning
Double DQNDecouples "choosing the action" from "evaluating the action" with two networks, relieving the systematic Q overestimation.Value Learning
Dueling DQNSplits Q into two output heads — state value V and advantage A — making value estimates more stable and learning faster.Value Learning
Prioritized ReplayWeights replay sampling by TD-error magnitude, so the samples that are "not yet learned" get replayed more often.Value Learning
RainbowCombines seven improvements — Double, Dueling, prioritized replay, multi-step returns, distributional Q (C51), noisy networks, and more — into a single agent.Value Learning
Policy GradientDifferentiates the policy parameters directly: raise the probability of high-return actions and push down low-return ones — naturally supports stochastic policies.Policy Gradient
REINFORCEThe plainest policy-gradient algorithm: uses the full-episode return as the weight — clear intuition, but enormous variance.Policy Gradient
BaselineA quantity subtracted from the return (usually the value function V); it leaves the gradient expectation unchanged while slashing variance.Policy Gradient
Actor-CriticA pairing of a policy network (actor) and a value network (critic); the critic supplies the actor with a low-variance advantage signal — the dominant architecture today.Actor-Critic Family
A2C / A3CActor-Critic running across parallel environments: A3C is the asynchronous version, A2C the synchronous one (more common and easier to reproduce).Actor-Critic Family
GAE (Generalized Advantage Estimation)Weights n-step advantages with a parameter λ, trading bias against variance along a continuous spectrum — PPO's standard companion.Actor-Critic Family
TRPOCaps each policy update with a KL-divergence constraint, guaranteeing monotonic improvement — but complex to implement and computationally heavy.Policy Gradient
PPOApproximates TRPO's constraint with a clipped objective — simple, stable, easy to implement, and the most mainstream on-policy algorithm today.Policy Gradient
DDPGDeterministic policy gradients + Q-learning: an off-policy continuous-control algorithm and the early form of "deep actor-critic".Actor-Critic Family
TD3Three fixes to DDPG: twin Q networks, delayed policy updates, and target policy smoothing — greatly relieving value overestimation and variance.Actor-Critic Family
SACA maximum-entropy off-policy algorithm whose objective additionally maximizes policy entropy — one of the de facto standards for continuous control.Actor-Critic Family
Entropy RegularizationAdds a policy-entropy term to the objective to encourage randomness, preventing premature convergence and preserving exploration.Exploration & Exploitation
Temperature Coefficient (α)In SAC, the coefficient weighting the entropy bonus; it can be tuned automatically (auto α adjustment), sparing you the manual search.Actor-Critic Family

A classic gotcha

SARSA vs. Q-learning: the two formulas differ by exactly one choice — SARSA uses the a′ actually taken, Q-learning uses the a′ picked by max. That single choice determines the on/off-policy property, and it comes up again and again in interviews and debugging. See Value Learning.

4. Deep RL Engineering Terms ​

High-frequency vocabulary for the "make it actually work" phase — mostly about why DQN/PPO work and how to deploy them.

TermOne-line definitionRelated page
Sample EfficiencyHow many environment interactions it takes to reach a given performance; off-policy algorithms (SAC, DQN) are typically far more sample-efficient than on-policy ones (PPO).Actor-Critic Family
Learning CurveA plot of return against training steps / environment interactions — the first diagnostic tool for RL tuning and evaluation.Evaluation & Benchmarks
SeedThe random seed; fixing it makes experiments reproducible, but tuning repeatedly against the same seed causes hidden overfitting.Evaluation & Benchmarks
ReproducibilityWhether the same code and config can reproduce similar results — requires fixed seeds, fully logged hyperparameters, and pinned environment versions.Common Pitfalls & Anti-Patterns
Distribution ShiftMismatch between training and deployment distributions (e.g., sim → real, offline → online); RL policies often collapse because of it.Offline RL
Domain RandomizationRandomizing simulated physics/rendering parameters during training so the policy learns invariances — a key enabler of zero-shot Sim2Real.Robotics Sim2Real
Sim2RealThe transfer from training in simulation to deploying in the real world; the core gap is modeling error in dynamics, observations, and contact.Robotics Sim2Real
Curriculum LearningA training schedule that gradually moves from easy tasks to hard ones, countering sparse rewards and exploration difficulty.Reward Engineering
HER (Hindsight Experience Replay)Rewrites "goals not achieved" as "achieved" to fabricate positive samples, solving sparse-reward and multi-goal problems.Reward Engineering
Intrinsic Reward / CuriosityUses intrinsic signals such as novelty or prediction error as extra rewards, driving the agent to explore unseen states.Exploration & Exploitation
Count-Based ExplorationCounts state visits and bonuses rare states — a naive method that works in tabular worlds and fails in high dimensions.Exploration & Exploitation
RND (Random Network Distillation)Uses a fixed random network's output as a "novelty target": states with large prediction error are more "novel" and earn higher intrinsic reward.Exploration & Exploitation
Vectorized EnvironmentsRunning multiple environment copies in parallel to collect experience — boosts throughput and training stability; standard engineering for A2C/PPO.Choosing Frameworks & Tools

5. Advanced Paradigm Terms ​

Vocabulary for frontier paradigms — offline RL, multi-agent, alignment — essential reading for 2020s papers.

TermOne-line definitionRelated page
Model-Based RLLearns an environment model first, then plans and learns inside it — sample-efficient, but model error compounds.Model-Based RL
World ModelA model that encodes state representations and predicts transitions and rewards — the core component of Dreamer, TD-MPC, and MuZero.Model-Based RL
Model-Free RLAlgorithms that skip the environment model and learn value/policy directly from interaction data — DQN, PPO, and SAC are all model-free.Value Learning
Offline RLTrains only on a fixed historical dataset, with no further environment interaction; the key difficulty is value overestimation on OOD actions.Offline RL
OOD (Out-of-Distribution)State-action pairs absent from, or barely covered by, the training data; in offline RL the "hallucinated overestimates" that Q networks extrapolate for them are the root disease.Offline RL
CQL (Conservative Q-Learning)Adds a penalty on the values of OOD actions to the Q-learning objective — an offline RL method that fights overestimation with "conservatism".Offline RL
IQL (Implicit Q-Learning)Regresses values only for in-distribution actions without extrapolation — an offline RL method that needs no explicit behavior regularizer.Offline RL
Behavior Cloning (BC)Direct supervised learning from expert demonstrations (state in, action out) — simple, but unable to handle states the expert never visited.RL vs. Neighboring Fields
Imitation LearningThe family of paradigms for learning policies from demonstrations: behavior cloning, inverse RL, GAIL, and more.RL vs. Neighboring Fields
Inverse RL (IRL)Infers the reward function from expert behavior — another path forward when reward engineering is too hard to write by hand.Reward Engineering
Multi-Agent RL (MARL)Multiple agents learn to make decisions simultaneously; the environment dynamics shift as others' policies change, making the problem non-stationary.Multi-Agent RL
CTDE"Centralized training, decentralized execution": global information is available during training, but deployment uses only local observations — the mainstream MARL paradigm.Multi-Agent RL
Nash EquilibriumA strategy profile where no agent can improve its payoff by unilaterally deviating — a convergence target for MARL that is often also where the difficulty lies.Multi-Agent RL
Non-StationarityThe environment dynamics keep changing because other agents are learning at the same time, so the single-agent MDP assumption breaks down.Multi-Agent RL
Value DecompositionDecomposes the joint Q value into a sum or product of individual Q values (e.g., QMIX) — the classic way to realize CTDE.Multi-Agent RL
MADDPGA multi-agent extension of DDPG with centralized critics and decentralized actors; off-policy, continuous actions.Multi-Agent RL
MAPPOThe multi-agent version of PPO; simple and direct, yet it has set new SOTA results on many tasks.Multi-Agent RL
RLHFA three-stage alignment pipeline: train a reward model from human preferences, then fine-tune the language model with PPO (SFT → RM → PPO).RLHF & Human-Feedback Alignment
Reward ModelA model trained on pairs of human preferences to score candidate text — it serves as "automated human feedback" in RLHF.RLHF & Human-Feedback Alignment
Bradley–Terry ModelThe standard preference-modeling assumption: P(choose A over B) = σ(r(A) − r(B)) — the foundation of reward-model training.RLHF & Human-Feedback Alignment
KL PenaltyA regularizer in RLHF that keeps the new policy from drifting too far from the reference model, preventing PPO from "flying off".RLHF & Human-Feedback Alignment
DPO (Direct Preference Optimization)An alignment shortcut that builds the policy objective directly from preference data, skipping both the reward model and PPO.RLHF & Human-Feedback Alignment
AlignmentThe overarching goal of making model behavior match human intent and values; RLHF, DPO, and RLAIF are all means to it.RLHF & Human-Feedback Alignment
Alignment TaxThe price of alignment training, paid in degraded general capability — the classic side effect of RLHF.RLHF & Human-Feedback Alignment
Reward Over-OptimizationThe phenomenon where the reward-model score climbs while true quality drops (Goodhart's law) — a typical failure mode of RLHF training.RLHF & Human-Feedback Alignment
RLAIFReplaces human feedback with AI feedback to train the reward model, scaling alignment up.RLHF & Human-Feedback Alignment

6. Easily Confused Terms ​

The five-plus confusion clusters below are the most common "danger zones" in interviews and paper reading, pulled out here for side-by-side comparison.

Confusable pairThe key distinctionOne-line memory hook
MC vs. TDMC uses full-episode returns (unbiased, high variance); TD uses "reward + bootstrapped estimate" (biased, low variance)"MC waits for the outcome; TD guesses one step"
On-policy vs. Off-policyWhether the data comes from the target policy; off-policy can reuse historical data"SARSA learns from itself; Q-learning learns from others"
Value Learning vs. Policy GradientValue learning learns Q first, then derives a policy; policy gradient learns the policy directly"Estimate-then-choose vs. choose directly"
Model-Based vs. Model-FreeWhether an explicit model of the transition/reward dynamics is learned"Do you need to learn the world first?"
RLHF vs. DPORLHF has three stages (RM + PPO); DPO is one-step, with no RM and no PPO"The detour vs. the direct route"
TRPO vs. PPOTRPO enforces a hard KL constraint; PPO replaces it with a soft clip approximation"A constraint turned into a penalty (approximation)"
DDPG vs. TD3 vs. SACAll three are off-policy continuous control; TD3 fixes DDPG's overestimation, and SAC adds maximum entropy on top"SAC = TD3's ideas + entropy"

7. Abbreviations A–Z ​

Every abbreviation on this page, alphabetized — when you can't remember where one came from, look it up here.

Abbrev.Full nameWhere on this page
A2C / A3CAdvantage Actor-Critic / Asynchronous A2CAlgorithm Families
ACActor-CriticAlgorithm Families
BCBehavior CloningAdvanced Paradigms
BTBradley–Terry preference modelAdvanced Paradigms
CQLConservative Q-LearningAdvanced Paradigms
CTDECentralized Training with Decentralized ExecutionAdvanced Paradigms
DDPGDeep Deterministic Policy GradientAlgorithm Families
DPODirect Preference OptimizationAdvanced Paradigms
DPDynamic ProgrammingAlgorithm Families
DQNDeep Q-NetworkAlgorithm Families
GAEGeneralized Advantage EstimationAlgorithm Families
HERHindsight Experience ReplayEngineering Terms
IQLImplicit Q-LearningAdvanced Paradigms
IRLInverse Reinforcement LearningAdvanced Paradigms
KLKullback–Leibler divergenceMath Primer
MARLMulti-Agent RLAdvanced Paradigms
MCMonte CarloAlgorithm Families
MDPMarkov Decision ProcessFoundations
OODOut-of-DistributionAdvanced Paradigms
POMDPPartially Observable MDPFoundations
PPOProximal Policy OptimizationAlgorithm Families
RMReward ModelAdvanced Paradigms
RNDRandom Network DistillationEngineering Terms
RLHFRL from Human FeedbackAdvanced Paradigms
SACSoft Actor-CriticAlgorithm Families
SARSAState-Action-Reward-State-ActionAlgorithm Families
SFTSupervised Fine-TuningRLHF
TDTemporal DifferenceAlgorithm Families
TD3Twin Delayed DDPGAlgorithm Families
TRPOTrust Region Policy OptimizationAlgorithm Families

DANGER

Abbreviations with the same spelling but different meanings are common in papers: TD can mean "temporal difference" in one paper and "task description" in another; in some contexts IML stands for "implicit imitation learning". Always pin down an abbreviation from its context, and in interviews, state the full name before elaborating.

Further Reading ​

References ​