Skip to content

Deep Reinforcement Learning

Quick overview Reinforcement learning enables agents to maximize cumulative rewards through trial and error; deep RL uses neural networks as its "brain." This article clarifies the MDP setup, the DQN family and Actor-Critic spectrum, the positioning of PPO/SAC, and extends to the exploration-exploitation trade-off, reward design, and the significance of RLHF for LLMs.

Deep Reinforcement Learning ​

One-sentence definition: Reinforcement learning (RL) studies how an agent learns optimal behavior policies through interaction with its environment, guided by "reward signals"; deep RL uses neural networks as approximation function for policies/value functions, enabling RL to move from tabular problems to high-dimensional continuous worlds like Go, games, and robot control.

1. Problem Setup: MDP, States, Actions, Rewards ​

The standard mathematical framework for RL is the Markov Decision Process (MDP), a five-tuple <S, A, P, R, γ>:

  • State S: the environment's observation at time t, s_t.
  • Action A: the action the agent can choose, a_t.
  • Transition probability P: P(s'|s, a), the environment's response to the new state.
  • Reward R: r_t, the immediate feedback for a single action.
  • Discount factor γ: the discount for future rewards, γ∈[0,1), depreciating "delayed rewards" — consistent with the intuition of "present over future."

The agent's goal: maximize cumulative (discounted) reward G_t = Σ_k γᵏ·r_{t+k}.

Three core concepts:

  • Policy π(a|s): the rule for choosing action a in state s (the policy network is its function approximation).
  • State value function V(s): the expected cumulative reward starting from s and following policy π.
  • Action value function Q(s,a): the expected cumulative reward by first taking a in s, then following π thereafter. The two satisfy the Bellman equation: V(s) = max_a [R(s,a) + γ·E[V(s')]] — recursively expressing "future value" as "current reward + next-state value." This is the mathematical foundation of virtually all RL algorithms.

2. RL in Relation to Supervised and Unsupervised Learning ​

RL sits in the "middle ground" between supervised and unsupervised learning, but is distinct from both:

  • Supervised learning has ground-truth answers (labels); RL only has "rewards" — sparse, delayed, possibly noisy signals, with no "ground truth" telling the agent which step was wrong.
  • Unsupervised learning discovers structure from data itself; RL's structure comes from "trajectories generated through interaction with the environment."
  • What's unique: RL's data is generated by the policy itself — the data distribution evolves as the policy evolves. This creates two major challenges: (1) samples are correlated and not i.i.d.; (2) old data becomes "stale" after the policy changes (distribution shift).

These are the deep reasons for RL's training instability and low sample efficiency. Its evaluation is also fundamentally different — there is no fixed test set; you must re-run the environment to sample. See Deep Learning Evaluation and Experiments and Deep RL Applications.

3. DQN and Its Improvements: Double, Dueling, PER ​

Deep Q-Network (DQN; Mnih et al., 2015) is a milestone in deep RL: using a neural network to fit the Q function and achieving superhuman performance on Atari games. It solved three technical challenges of "using deep networks for RL":

  1. Replay buffer: store past transitions (s,a,r,s') in a buffer, randomly sample during training — break sample correlations, improve sample reuse (this is the remedy for "broken i.i.d.").
  2. Target network: use a delayed-updated parameter copy to compute target values r + γ·max Q(s',a'), avoiding divergence from "chasing yourself."
  3. Stability of ∂Q/∂θ: the loss (r + γ·max Q(s',a') − Q(s,a))², paired with Adam and gradient clipping from Optimization and Gradient Descent.

Three classic improvements in the DQN family:

  • Double DQN (2016): decouple the max operation by "using the current network to select actions and the target network to provide values," eliminating Q-value overestimation.
  • Dueling DQN (2016): split Q into V(s) + A(s,a) (state value + advantage), letting the network learn "which states are inherently valuable," improving generalization.
  • PER (Prioritized Experience Replay, 2016): weight-sample transitions by TD error (the gap between prediction and target) — train more on "highly wrong" samples, improving sample efficiency.

4. Policy Gradient and Actor-Critic ​

Value-based methods (DQN family) learn Q and derive policies from it; policy gradient directly parameterizes the policy and updates along the "gradient of expected reward":

∇θ J(θ) = E[ ∇θ log π(a|s) · A(s,a) ]

Intuition: increase the probability of actions that "do better than average," decrease those that do worse (A is the advantage, see below). Policy gradients natively support continuous action spaces (DQN can't handle continuous actions), which is why they dominate in robot control.

Actor-Critic combines the two:

  • Actor (policy): responsible for choosing actions, updating along the advantage.
  • Critic (value network): responsible for "scoring" actions (estimating advantage A = Q − V) for the Actor to use.

The division of labor is like "rider + navigator": the navigator provides finer-grained feedback than raw rewards (the advantage), and the rider adjusts the steering accordingly. This family includes A2C/A3C, DDPG, TD3, and more.

5. PPO and SAC: Two Main Algorithms ​

PPO (Proximal Policy Optimization; Schulman et al., 2017) is currently the most widely applied RL algorithm. It solves the problem of policy gradient "taking too big a step and falling": it uses the ratio of old/new policy probabilities r(θ) = π_θ/π_old with clipping to bound the magnitude of each update within a trusted region — "improve only a little each time, don't try to reach the sky in one leap."

L = E[ min(r·A, clip(r, 1−ε, 1+ε)·A) ]

PPO's strengths: simple implementation, relatively robust hyperparameters, decent sample efficiency. Many results from OpenAI/DeepMind (Dota, the RLHF stage of ChatGPT) use it.

SAC (Soft Actor-Critic; Haarnoja et al., 2018) is the other pole in continuous control: it adds an entropy maximization term to the objective (encouraging randomness, proactive exploration), paired with "soft updates" and dual Q networks, delivering excellent stability and sample efficiency. Rule of thumb: discrete/large-scale scenarios: PPO first; continuous control: consider SAC.

6. The Exploration-Exploitation Trade-off ​

The core tension in RL: exploration (trying new actions to gather information) vs. exploitation (acting on known optima). Exploit only, and you're trapped in local optima; explore only, and you'll never converge to the optimum.

Classic mechanisms:

  • ε-greedy: take random actions with probability ε (standard for DQN) — simple but directionless exploration.
  • UCB / confidence bounds: preferentially try actions with "high uncertainty."
  • Entropy regularization: keep the policy stochastic (SAC's entropy term, PPO's entropy bonus).
  • Intrinsic reward (curiosity): give extra rewards for "poorly predicted / novel" states, encouraging exploration of unknown areas.

7. Sample Efficiency and Reward Design ​

Two real-world pain points in RL:

  1. Sample efficiency: training an Atari agent requires tens of millions of environment steps; for real robots, each step costs real money. Remedies: replay buffers, model-predictive control (world models), Sim-to-Real transfer (train in simulation, fine-tune in the real world).
  2. Reward design: the reward signal is the "learning objective." Poor design teaches opportunistic behavior (reward hacking) — e.g., a cleaning robot learning to "hide the trash" rather than "clean it up." Principle: rewards should be sparse but semantically correct — better to give fewer signals than wrong ones. When using reward shaping, be careful not to change the optimal policy.

Engineering of data and environments (simulators, trajectory datasets) is covered in Data and Data Engineering.

8. RLHF: RL Enters LLMs ​

RLHF (Reinforcement Learning from Human Feedback) is RL's most important application in LLMs and a key factor in ChatGPT's success, in three steps:

  1. Collect preferences: have humans rank (or score) two responses to the same prompt.
  2. Train a reward model (RM): train a model that "predicts which response is better" using preference data (it's the ranking/contrastive loss from Loss Functions and Output Layers).
  3. RL fine-tuning: use PPO to have the LLM generate responses that the RM scores highly — maximizing RM score + KL penalty (the KL penalty prevents the model from straying too far from the SFT model and talking nonsense).

Why RLHF when LLMs can already "predict the next token"? Because "high next-token probability" ≠ "aligns with human preferences" (useful, harmless, honest). RLHF uses human preferences as the ultimate reward, directly aligning model behavior. See Large Language Models (LLM) for the full mechanism. Subsequent methods like DPO (Direct Preference Optimization) bypass the reward model and PPO, achieving alignment in a simpler way.

9. Trade-offs ​

Trade-offs

Value methods vs. policy gradients: the DQN family has high sample efficiency and works well for discrete actions, but has weak expressive power for continuous actions and stochastic policies; policy gradients natively handle continuous and stochastic policies, but suffer from high variance and low sample efficiency. Actor-Critic is the compromise between the two; PPO/SAC are the current engineering answers to that compromise.

Sample efficiency vs. implementation complexity: SAC's entropy regularization improves sample efficiency, but the entropy coefficient needs tuning; PPO is simple and robust but has average sample efficiency. Cheap data (simulation) → go PPO; expensive data (real environment) → prioritize SAC/world models.

Exploration randomness vs. convergence stability: entropy/intrinsic rewards promote exploration but slow convergence; insufficient exploration leads to suboptimal solutions. In practice, use "annealing" — explore more early in training, narrow down later.

Generality vs. task specificity: RL algorithms are sensitive to each environment's rewards and state encoding (change the reward scaling and hyperparameters collapse) — far from the "plug-and-play" nature of supervised learning. RL projects must budget heavily for tuning; see Training Recipes and Hyperparameter Tuning for recipes.

Deep reinforcement learning turns "learning to make decisions from experience" into engineering reality: from games to robotics to alignment in Large Language Models (LLM). Its foundations remain the networks from Neural Network Fundamentals and optimizers from Optimization and Gradient Descent — only the "data" is replaced by trajectories produced by the policy itself, which is why every component (replay buffer, target network, PPO clipping) radiates the wisdom of "how to tame non-i.i.d. data." See Deep RL Applications for the full application landscape.

Further Reading ​

References ​