Appearance
Policy Gradient Methods
In one sentence: this page covers policy gradient methods — instead of learning "how much each action is worth," you learn a policy network that directly encodes "what to do." We'll go from REINFORCE's high variance, through baselines and advantage, to PPO's clipped objective and GAE. By the end, you'll be able to explain the intuition behind the policy gradient theorem, write the PPO loss correctly, and know how to parameterize distributions for continuous action spaces.
1. Why Learn the Policy Directly: Three Shortcomings of Value-Based Learning
Value-based learning (learn Q → derive the policy via argmax) has three inherent limitations:
| Shortcoming | The value-based problem | The policy gradient fix |
|---|---|---|
| Continuous actions | argmax Q can't be computed in continuous spaces | The policy network directly outputs a continuous action distribution |
| Stochastic optimal policies | Q only yields deterministic actions | Policy gradients output distributions natively (poker and games need randomness) |
| High-dimensional / redundant states | Every state needs an accurate Q estimate | The policy only cares about "what to do" — naturally more economical |
| Constrained action spaces | Hard to express "action X is forbidden" | Can be modeled at the distribution level (masks, constraints) |
Value-based learning fits scenarios with "discrete actions where you can afford to evaluate Q for each one"; policy gradients are more general and form the backbone of modern mainstream algorithms like PPO and SAC. The two approaches eventually converge in Actor-Critic (see the Actor-Critic family).
An intuitive analogy
Value-based learning is like "scoring every dish on the menu and ordering the highest-rated one"; policy gradient is like "learning the craft of what to order in each situation." The former requires being able to enumerate every dish — the latter doesn't.
2. The Policy Network and the Objective
1. The Policy Network π_θ(a|s)
A neural network with parameters $\theta$ represents the policy: it takes state $s$ as input and outputs an action distribution. Two action-space flavors:
text
Discrete actions: output a softmax probability vector
[π(a1|s), π(a2|s), ..., π(aK|s)]
Continuous actions: output Gaussian parameters (mean + variance)
mean μ_θ(s), variance σ_θ(s) → a ~ N(μ, σ²)2. The Objective: Maximize Expected Return
Policy gradient methods optimize the expected discounted return directly:
$$ J(\theta) = \mathbb{E}{\tau \sim \pi\theta} \left[ \sum_{t=0}^{T} \gamma^t r_t \right] = \mathbb{E}{\tau \sim \pi\theta}[R(\tau)] $$
where the trajectory is $\tau = (s_0, a_0, r_1, s_1, \dots)$. Note carefully: the expectation is taken over trajectories sampled from π_θ itself — the objective's value depends on how the samples were collected. This is the most fundamental difference from supervised learning, and it is also the root of the method's high variance.
3. The Policy Gradient Theorem: Intuition First
1. The Plainest Intuition
We want $J(\theta)$ to go up, so we follow the gradient $\nabla_\theta J(\theta)$. The catch: the expectation in the objective depends on θ itself (changing θ changes the sampling distribution), so we can't differentiate through it directly. The policy gradient theorem provides an elegant rewrite:
$$ \nabla_\theta J(\theta) = \mathbb{E}{\tau \sim \pi\theta} \left[ \sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) , R(\tau) \right] $$
This is the key formula (REINFORCE uses exactly this form). Term by term:
- $\nabla_\theta \log \pi_\theta(a_t|s_t)$: the direction that increases the action's probability — every increase in probability nudges the parameters along this direction;
- $R(\tau)$: the total return of the entire trajectory, acting as the weight for "is this trajectory worth repeating."
Put them together: for high-return trajectories, we raise the probability of the actions that produced them; for low-return trajectories, we push those probabilities down. That is the entire intuition of policy gradients — "increase the probability of high-return actions, decrease the probability of low-return actions."
2. Where Does the Log Come From?
The $\nabla_\theta \log \pi_\theta$ comes from the "score function trick" (a.k.a. the REINFORCE trick):
$$ \nabla_\theta \mathbb{E}{\pi\theta}[R] = \mathbb{E}{\pi\theta}[R \cdot \nabla_\theta \log \pi_\theta] $$
It turns "differentiating through a distribution" into "differentiating the log-probability and multiplying by the return," which makes Monte Carlo estimation from samples possible. This is the mathematical foundation of the entire policy gradient family.
A mnemonic analogy
Policy gradient = "basketball shooting practice": keep rehearsing the posture that sinks the shot (high return → probability up), and drop the posture that misses (low return → probability down). Note that it never asks "how many points was this shot worth" — only "did we win the game or not" ($R(\tau)$ is the return of the whole episode).
4. REINFORCE: The Plainest Policy Gradient
1. The Algorithm
REINFORCE (Williams, 1992) updates using the return of the entire trajectory:
python
def reinforce(env, policy, episodes=1000, gamma=0.99, lr=0.01):
for _ in range(episodes):
# 1. Sample a full trajectory with the current policy
trajectory = collect_episode(env, policy) # [(s, a, r), ...]
# 2. Compute the discounted return G_t for each step, backwards
G = 0.0
for t, (s, a, r) in reversed(list(enumerate(trajectory))):
G = r + gamma * G # return G_t
# 3. Update along the "log-prob × return" gradient
policy.update(s, a, lr * G * grad_log_prob(s, a))2. The High-Variance Problem: REINFORCE's Fundamental Flaw
REINFORCE weights everything by the full-episode return $R(\tau)$. The problem:
- $R(\tau)$ mixes "the policy is good" with "the environment was lucky": the same policy can return 100 in a lucky episode and 0 in an unlucky one;
- so the gradient direction is dominated by luck, and the variance explodes;
- high-variance estimates need enormous amounts of data to converge — REINFORCE barely moves on anything beyond toy tasks.
Feel the high variance yourself
In version two of Progressive Tutorial: Getting Three Versions of Gymnasium Running, you'll implement REINFORCE yourself and watch its learning curve: either violent oscillation or no improvement at all. That's not a bug — it's the innate variance of policy gradients.
5. Baselines and Advantage: De-Meaning the Return
1. Intuition: You Need a Reference Point to Judge Good or Bad
"Is a return of 10 good or bad?" — it depends on the baseline. If all trajectories average 50, then 10 is bad; if they average -5, then 10 is excellent. REINFORCE weights by the absolute return and ignores this frame of reference entirely.
The baseline trick: subtract a baseline $b(s_t)$ that doesn't depend on the current action:
$$ \nabla_\theta J(\theta) = \mathbb{E}\left[ \sum_t \nabla_\theta \log \pi_\theta(a_t|s_t) , \big( R(\tau) - b(s_t) \big) \right] $$
Because $\mathbb{E}[\nabla_\theta \log \pi_\theta \cdot b(s_t)] = 0$ (provable: the baseline's expected contribution to the gradient is zero), subtracting it leaves the expected gradient unchanged (unbiased) and only reduces the variance.
2. The Advantage Function: The Most Natural Baseline
The best baseline is the state value function $V^\pi(s_t)$. That gives us the advantage function:
$$ A^\pi(s_t, a_t) = Q^\pi(s_t, a_t) - V^\pi(s_t) $$
Intuitively, $A$ answers "how much better is this action than average" — Q is this action's expected return, V is the average return over all actions, and a positive difference means the action beats the average. Replacing the total return with advantage, the gradient becomes:
$$ \nabla_\theta J(\theta) = \mathbb{E}\left[ \sum_t \nabla_\theta \log \pi_\theta(a_t|s_t) , A(s_t, a_t) \right] $$
3. How Do We Estimate A? — We Need a Second Network
Problem: $A = Q - V$, and we don't even have Q. Two engineering solutions:
- Option one (Actor-Critic): learn a value network $V_\phi(s)$ as the baseline and approximate the advantage with a "one-step return":
$$ \hat A_t = r_{t+1} + \gamma V_\phi(s_{t+1}) - V_\phi(s_t) $$
This has exactly the form of the TD error (see the TD section of value-based learning)! The combination of an Actor (policy network) + Critic (value network) is the Actor-Critic architecture — see the Actor-Critic family for details.
- Option two (GAE): a weighted combination of multi-step advantages — covered in Section 6 below.
6. GAE: Generalized Advantage Estimation
1. Starting from the TD Error
Define the one-step TD error:
$$ \delta_t = r_{t+1} + \gamma V(s_{t+1}) - V(s_t) $$
GAE (Generalized Advantage Estimation, Schulman et al., 2016) combines n-step advantages with exponentially decaying weights:
$$ \hat A^{GAE(\gamma,\lambda)}t = \sum^{\infty} (\gamma\lambda)^l , \delta_{t+l} $$
2. Intuition: λ Is a Bias-Variance Dial
| Value of λ | Equivalent to | Bias | Variance |
|---|---|---|---|
| λ=0 | 1-step TD only (heavy bootstrapping) | High | Low |
| λ=1 | Full trajectory (Monte Carlo style) | Low (unbiased) | High |
| λ=0.95 | The compromise (common in PPO/SAC) | Medium | Medium |
Intuition: λ controls "how many steps into the future do we trust" — a small λ (0) trusts only one step of experience (fast but biased), while a large λ (1) trusts all the way to the end of the trajectory (slow but unbiased). GAE is a continuous interpolation between MC and TD.
Practical defaults
In PPO/SAC, λ is usually set between 0.95 and 0.99. λ is the direct dial for "more variance vs. less bias," and it's one of the first three knobs you should touch when tuning.
7. TRPO: Why We Need a Trust Region
1. The Problem: Large Steps Collapse the Policy
Plain policy gradients update the parameters with a fixed learning rate. But distance in parameter space ≠ distance between policies: a parameter nudge of 0.01 can turn the policy distribution upside down (softmax is notoriously sensitive to its parameters in high dimensions). One oversized update → the policy collapses in an instant → and it may never recover. This is the usual story behind "my policy gradient training blew up."
2. TRPO's Answer: Constrain the Step Size
TRPO (Trust Region Policy Optimization, Schulman et al., 2015) confines each update to a trust region — it uses KL divergence to keep the new policy from straying too far from the old one:
$$ \max_\theta \mathbb{E}\left[ \frac{\pi_\theta(a|s)}{\pi_{\theta_{old}}(a|s)} \hat A_t \right] \quad \text{s.t.} \quad \mathbb{E}\left[ D_{KL}(\pi_{\theta_{old}} | \pi_\theta) \right] \le \delta $$
Intuition: every update guarantees "the KL distance between the new and old policies stays within δ," so you can never wander off in one giant step. TRPO is beautiful in theory but painful to implement (it needs conjugate gradients to solve the constrained problem).
3. TRPO's Successor: PPO
PPO (Proximal Policy Optimization, Schulman et al., 2017) approximates TRPO's trust-region effect with clipping, is vastly simpler to implement, and has become the de facto industry standard.
8. PPO's Clipped Objective: Formula and Implementation
1. The Core Objective
Define the probability ratio (importance sampling ratio):
$$ \rho_t(\theta) = \frac{\pi_\theta(a_t | s_t)}{\pi_{\theta_{old}}(a_t | s_t)} $$
PPO's clipped objective:
$$ L^{CLIP}(\theta) = \mathbb{E}_t \left[ \min\left( \rho_t(\theta) \hat A_t, ; \operatorname{clip}(\rho_t(\theta), 1-\epsilon, 1+\epsilon) , \hat A_t \right) \right] $$
- $\rho_t \hat A_t$: the plain policy gradient term (ratio × advantage);
- $\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)$: clips the ratio to $[1-\epsilon, 1+\epsilon]$;
- $\min$: takes the smaller of the two.
2. Why Clipping Works: A Case-by-Case Breakdown
There are four cases (assume ε=0.2; when $\hat A_t>0$ we want to raise this action's probability, when $\hat A_t<0$ we want to lower it):
| Case | What ρ means | Effect of clip |
|---|---|---|
| A>0, probability rising (ρ>1) | The new policy likes this action more | ρ is clamped at 1.2 — don't get greedy |
| A>0, probability falling (ρ<1) | The new policy likes it less | The gradient encourages an increase; no clipping needed (we want it to rise anyway) |
| A<0, probability rising (ρ>1) | The action is actually bad | No gradient penalty (min takes the clipped side) — neither encouraged nor forced |
| A<0, probability falling (ρ<1) | The right direction | Clamped at 0.8 — limits how much a single step can change |
In one line: "don't get carried away with good news, don't panic over bad news" — PPO is happy to push bad actions down and good actions up, but no single step is allowed to overdo it (gradients are cut off beyond ±ε). This replaces TRPO's "global KL constraint" with "per-sample clipping": simple, and stable.
3. PPO's Full Loss: A Sum of Three Terms
$$ L^{PPO}(\theta) = \underbrace{L^{CLIP}}{\text{policy}} - c_1 \underbrace{L^{VF}}{\text{value}} + c_2 \underbrace{\mathcal{H}[\pi_\theta]}_{\text{entropy reg.}} $$
- $L^{VF}$: the Critic's (value network's) regression loss $(V_\phi(s_t) - \hat R_t)^2$;
- $\mathcal{H}$: the policy entropy term, with coefficient $c_2$ controlling the strength of exploration (see entropy regularization in Exploration and Exploitation).
PPO has two variants: the clipped version (described above, most common) and the KL-penalty version (KL divergence as a penalty term — a soft constraint à la TRPO). In practice, the clipped version is the default.
4. The PPO Training Loop (Pseudocode)
python
def ppo_iteration(env, actor, critic, n_steps=2048, epochs=10, eps_clip=0.2):
# 1. Sample a batch of trajectories with the old policy (rollout)
batch = collect_rollouts(env, actor, n_steps) # n_steps transitions of (s,a,r,s',done)
# 2. Compute advantages (GAE)
adv = compute_gae(batch, critic, gamma=0.99, lam=0.95)
# 3. Run several epochs of minibatch gradient ascent on the batch
for _ in range(epochs):
for mb in sample_minibatches(batch):
ratio = actor.prob(mb.a) / actor_old.prob(mb.a) # ρ
loss = -torch.min(ratio * mb.adv,
torch.clamp(ratio, 1-eps_clip, 1+eps_clip) * mb.adv)
loss += -0.01 * actor.entropy(mb.s) # entropy regularization
loss += critic_loss(mb) # value network
optimizer.step(loss)
# 4. Copy the actor's parameters into actor_old for the next roundKey point: multiple epochs of minibatch updates (far more data-efficient than REINFORCE's single update per batch) is one of the engineering keys behind PPO's stability and speed.
Common PPO pitfalls
- The clip threshold ε defaults to 0.2, but it's sensitive to reward scale: with huge rewards, A becomes huge and clipping can't hold the line — normalize advantages first (subtract the mean, divide by the standard deviation);
- Don't skip tuning the entropy coefficient c_2: too large and the policy stays random forever, too small and it collapses prematurely — sweep somewhere between 0 and 0.01;
- Think about GAE's λ and γ together: γ decides "how far you look ahead," λ decides "how many steps of experience you trust";
- Handling
done(no bootstrapping at terminal states) is just as easy to get wrong in PPO.
9. Continuous Action Spaces: Distribution Choice and Implementation Notes
1. Gaussian Policies: The Default Choice
Continuous actions default to a diagonal Gaussian: the network outputs a mean $\mu_\theta(s)$ and a standard deviation $\sigma_\theta(s)$ (or log-standard-deviation), and we sample $a \sim \mathcal{N}(\mu, \sigma^2 I)$.
text
Continuous policy network output (two heads):
┌───────────┐ ┌─────┐
│ state s │ ─────▶ │ μ(s)│ → action mean (tanh-squashed to the action range)
│ input │ ├─────┤
│ │ ─────▶ │ σ(s)│ → exploration variance (learnable or fixed)
└───────────┘ └─────┘2. Key Implementation Details
| Detail | Why it matters |
|---|---|
| Initial value of log std | Too small → too little exploration early on; usually initialize to -1 ~ -2 (mild exploration) |
| tanh squashing | Maps Gaussian samples to the action bounds [-1,1], but you must correct the log-probability with the Jacobian |
| Entropy computation | Gaussian entropy has the closed form $\frac{1}{2}\ln(2\pi e \sigma^2)$ — use it to tune exploration via c_2 |
| Learnable vs. fixed variance | Learnable variance is unstable but adapts; fixed variance is simple but needs hand-tuning |
3. Discrete Actions and Hybrid Spaces
- Discrete: softmax outputs class probabilities (Categorical);
- Hybrid (discrete + continuous, e.g., "choose a tool + move in a direction"): split into multiple distribution heads and sum the log-probabilities.
10. Practical Recommendations
| Scenario | Recommendation | Rationale |
|---|---|---|
| Learning the basics | REINFORCE (tutorial version) | The shortest path to seeing the gradient's essence; see Progressive Tutorial |
| General discrete/continuous tasks | PPO | Stable, widely implemented, few hyperparameters |
| Sample efficiency first | SAC (see the Actor-Critic family) | Off-policy; 10–100× fewer samples |
| Robotics Sim2Real | PPO/SAC + domain randomization | See Robotics Control and Sim2Real |
| Large-scale parallel training | PPO + many environments | PPO supports parallel rollouts natively |
A frequent interview question — "Why did PPO replace TRPO?"
TRPO requires solving a constrained quadratic program (via conjugate gradients) — engineering-heavy and computationally expensive. PPO achieves a similar "trust region" effect with first-order clipping: ~20 lines of code, fast and stable. The engineering world chose simplicity.
Further Reading
- The Actor-Critic Family — where policy networks and value networks converge: A2C, DDPG, TD3, SAC
- Value-Based Learning: From Dynamic Programming to DQN — the TD error of the value network is precisely the raw material for advantage estimation
- Exploration and Exploitation — built-in exploration for the policy gradient family: entropy regularization, parameter noise
- Robotics Control and Sim2Real — PPO/SAC in the starring role on real robots
- Classic Paper Deep Dives — a paragraph-by-paragraph close reading of the PPO paper, with interview Q&A
- Progressive Tutorial: Getting Three Versions of Gymnasium Running — hand-written implementations from REINFORCE to PPO
References
- Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3-4), 229-256. The original REINFORCE paper.
- Sutton, R. S., McAllester, D., Singh, S., & Mansour, Y. (1999). Policy Gradient Methods for Reinforcement Learning with Function Approximation. NeurIPS. The proof of the policy gradient theorem.
- Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust Region Policy Optimization. ICML. arXiv:1502.05477
- Schulman, J., Wolski, P., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347
- Schulman, J., Moritz, P., Levine, S., Jordan, M., & Abbeel, P. (2016). High-Dimensional Continuous Control Using Generalized Advantage Estimation. ICLR. arXiv:1506.02438
- OpenAI Spinning Up. Intro to Policy Optimization. https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html