Appearance
The Actor-Critic Family
In one sentence: this page covers the Actor-Critic family — how the "actor" (policy network) and the "critic" (value network) merge into the mainstream architecture of modern RL. By the end, you'll be able to draw the family tree of A2C, DDPG, TD3, SAC, and PPO, understand where on-policy and off-policy implementations diverge, and pick the right algorithm for a given task.
1. The Actor-Critic Architecture: An Actor and a Judge
1. Why Two Networks?
Recall policy gradient methods: REINFORCE weights the gradient by the full-episode return $R(\tau)$, which has enormous variance. The fix is to replace the total return with the advantage $A(s,a) = Q(s,a) - V(s)$ — and estimating the advantage requires a value network.
So the two networks have clearly divided roles:
text
┌────────────────────────────────────────┐
│ Actor-Critic │
│ │
│ ┌───────────┐ ┌───────────┐ │
│ │ Actor │ │ Critic │ │
│ │ policy │ │ value │ │
│ │ net π_θ │ │ net V_φ │ │
│ └─────┬─────┘ └─────┬─────┘ │
│ │ outputs actions │ judges │
│ ▼ ▼ │
│ interacts ──▶ (s,a,r,s') ──▶ TD │
│ with env error / adv. │
│ │
│ Actor goal: make high-advantage │
│ actions more likely │
│ Critic goal: make V accurate │
│ (TD regression) │
└────────────────────────────────────────┘- Actor: learns the policy $\pi_\theta(a|s)$; its goal is to maximize the advantage-weighted log-probability (see Sections 5 and 6 of policy gradient methods);
- Critic: learns the value $V_\phi(s)$, scores "how good is this situation," and supplies the advantage signal to the Actor — while improving itself at the same time (TD regression).
Why the Critic reduces variance
The Critic plays the "baseline" role: $A = Q - V$ strips "how good the action itself is" out of "how good the situation itself is," and what remains has much lower variance. The Actor only cares about "relative to average, what makes this action good."
2. How Actor and Critic Update Together (A2C as an Example)
python
def a2c_update(actor, critic, transitions):
# transitions: a batch of (s, a, r, s', done)
# 1. The Critic provides the baseline + computes the TD target
v = critic(s)
v_next = critic(s_next)
td_target = r + gamma * v_next * (1 - done) # no bootstrapping when done
advantage = td_target - v # use the TD error as the advantage
# 2. Critic update: push V toward the TD target
critic_loss = (td_target - v).square().mean()
# 3. Actor update: raise the probability of high-advantage actions
actor_loss = -(log_prob(a) * advantage.detach()).mean()
# (+ entropy regularization term, see the exploration page)Note that detach() cuts the gradient through advantage — during the Actor update, the advantage is treated as a constant and no gradient flows through it.
2. A2C/A3C: Parallel Environments and Advantage
1. Background: Samples from a Single Environment Are Slow and Correlated
Policy gradients are on-policy: data is used once and thrown away, so you must keep collecting fresh data. Serial sampling from a single environment is painfully slow, and consecutive samples are strongly correlated. A3C (Asynchronous Advantage Actor-Critic, Mnih et al., 2016) throws multiple parallel workers at their own environments:
- Each worker collects data independently, computes gradients independently, and pushes them asynchronously to shared parameters;
- A2C (the synchronous version): drop the asynchrony — all workers collect one batch in lockstep and update once together.
| Dimension | A3C (asynchronous) | A2C (synchronous) |
|---|---|---|
| How updates happen | Each worker pushes gradients asynchronously | Data from all workers is collected and updated in one go |
| Implementation complexity | High (locks, communication, parameter races) | Low (a single batched update) |
| Track record | The historical version | Usually on par or better (A2C is the default now) |
A2C commonly estimates advantage with n-step returns: let each worker run for n steps, then compute everything at once:
$$ \hat A_t = \sum_{k=0}^{n-1} \gamma^k r_{t+k} + \gamma^n V(s_{t+n}) - V(s_t) $$
This is a discrete approximation of GAE (an n-step truncation at λ=1; note that at λ=0, GAE degenerates to the 1-step TD error — not an n-step return). For the theory, see Section 6 of policy gradient methods.
2. A2C in One Sentence
A2C = policy gradient + value baseline + multiple parallel environments. It's the most "textbook" member of the Actor-Critic family, and while PPO has largely replaced it, its parallelization idea (vectorized environments) is the infrastructure for every later algorithm.
3. DDPG: Deterministic Policy Gradients for Continuous Actions
1. The Problem: Softmax/Gaussian Doesn't Cut It in Continuous Action Spaces
Everything so far has been a stochastic policy $\pi(a|s)$. But in continuous control (robot joint torques), sampling noise from a stochastic policy is a nuisance, and the importance ratio has enormous variance under off-policy learning. DDPG (Deep Deterministic Policy Gradient, Lillicrap et al., 2016) takes a different route: a deterministic policy $\mu_\theta(s)$ (state in, action out directly), paired with a Critic that learns $Q$.
2. DPG Intuition: Use the Critic's Gradient to "Coach" the Action
The deterministic policy gradient points the Actor's output action in the direction where Q increases:
$$ \nabla_\theta J \approx \mathbb{E}\left[ \nabla_a Q(s,a) \big|{a=\mu\theta(s)} \cdot \nabla_\theta \mu_\theta(s) \right] $$
Intuition: the Critic says "nudge the action this way and Q will go up," and the Actor complies. The Critic becomes the Actor's "action coach" — instead of log-probabilities and advantage as in stochastic policy gradients, we differentiate Q directly with respect to the action.
3. Four Components (the Off-Policy Standard Kit)
A complete DDPG implementation needs four ingredients (later inherited by TD3 and SAC):
| Component | Role |
|---|---|
| Online Actor network μ_θ | Outputs deterministic actions |
| Online Critic network Q_φ | Evaluates the value of (s,a) |
| Replay buffer | Reuses samples off-policy (see experience replay in value-based learning) |
| Target network (soft update) | Slowly tracks the online parameters with $\tau$: $\phi' \leftarrow \tau\phi + (1-\tau)\phi'$ |
DDPG's notorious instability
DDPG is notoriously hard to tune: Q overestimation (the max operator amplified by the deterministic policy), extreme sensitivity to hyperparameters, and hard-to-tune exploration (Ornstein-Uhlenbeck noise). It serves as the "prototype before TD3" — in practice, go straight to TD3 or SAC.
4. TD3: Three Cures for DDPG's Three Ailments
TD3 (Twin Delayed DDPG, Fujimoto et al., 2018) prescribes three remedies for DDPG's three problems:
| Ailment | Remedy |
|---|---|
| Q overestimation (max noise accumulates) | Twin Q networks: two Critics, take the smaller value $Q = \min(Q_1, Q_2)$ (clipped double Q-learning) |
| Actor updates too frequently | Delayed updates: the Actor updates once for every two Critic updates |
| Deterministic policy overfitting | Target policy smoothing: add noise to target actions $\tilde a = \mu(s') + \epsilon$, $\epsilon \sim \text{clip}(\mathcal{N}(0,\sigma^2), -c, c)$ |
Intuition: taking the min of the twin Q networks breaks the path where "max turns noise into signal" (the continuous-action cousin of Double DQN); delayed updates let the Critic become accurate before its signal feeds the Actor; and the smoothing noise keeps the value function smooth for similar actions around the policy.
5. SAC: Maximum-Entropy RL
1. Core Idea: The Objective Is Not Just "High Return" but Also "High Action Entropy"
SAC (Soft Actor-Critic, Haarnoja et al., 2018) rewrites the objective in maximum-entropy form:
$$ J(\pi) = \mathbb{E}\left[ \sum_t \gamma^t \left( r_t + \alpha , \mathcal{H}(\pi(\cdot | s_t)) \right) \right] $$
The entropy term $\mathcal{H}(\pi(\cdot|s))$ measures "how random is the policy in this state" (see Exploration and Exploitation). SAC writes "keep exploring" explicitly into the objective — "high return and diverse actions, both."
Intuitively, this behaves like "be as random as possible as long as the return barely suffers." The payoff:
- Exploration and robustness: high entropy → built-in exploration, natural resistance to overfitting, and insensitivity to reward noise;
- Multi-modal tasks: some problems have several equally good solutions (detouring left or right, say), and SAC keeps all of them alive instead of locking into one.
2. The Soft Bellman Equation and Soft Q
SAC's Critic learns "soft" values, folding entropy into Q:
$$ Q(s,a) = r + \gamma , \mathbb{E}\left[ Q(s', a') - \alpha \log \pi(a'|s') \right] $$
Note the target no longer uses $\max Q$ but "expected Q minus the entropy penalty." SAC is an off-policy stochastic-policy algorithm (Q-learning-style updates + Gaussian policy), which gives it both the sample efficiency of off-policy learning and the exploration ability of a stochastic policy — the fundamental reason it wins across the board in continuous control.
3. Automatic Temperature α
The entropy coefficient α decides "how much we care about exploration": too large → the policy is too random (ignores the task); too small → SAC degenerates into plain RL. SAC uses automatic temperature tuning: set a target entropy (usually $\mathcal{H}_{target} = -\dim(\mathcal{A})$) and adjust α during training so the actual entropy approaches the target:
$$ \min_\alpha \mathbb{E}\left[ -\alpha \log \pi(a|s) - \alpha \mathcal{H}_{target} \right] $$
Intuition: policy too deterministic (entropy too low) → raise α to force more randomness; too random → lower α to refocus on the task. This one mechanism saves engineers the single most sensitive exploration hyperparameter.
4. Why Continuous Control Is SAC's Home Turf
| Property | Why it matters for continuous control |
|---|---|
| Off-policy (replay reuse) | Real-robot / expensive-simulator samples are costly; sample efficiency is critical |
| Stochastic policy (entropy) | Continuous actions need persistent exploration; deterministic-policy noise is hard to tune |
| Twin Q + target smoothing (inherited from TD3) | Cures Q overestimation; stable training |
| Automatic α | Eliminates the exploration-hyperparameter search |
A head-to-head comparison of SAC versus its main rivals:
| Dimension | SAC | TD3 | PPO |
|---|---|---|---|
| On/off-policy | Off | Off | On |
| Policy type | Stochastic (Gaussian) | Deterministic | Stochastic (Gaussian / categorical) |
| Sample efficiency | Highest | High | Low (10–100× worse) |
| Training stability | Good (twin Q + automatic α) | Good | Very good |
| Hyperparameter sensitivity | Low | Medium | Medium |
| Parallelizability | Mediocre | Mediocre | Excellent (parallel rollouts are natural) |
| Recommended for | Costly samples, continuous control | Continuous control where deterministic is fine | General use, large-scale parallelism |
6. PPO as a Special Case of Actor-Critic
PPO is covered in detail on the policy gradient methods page. In the Actor-Critic family tree, it is simply an on-policy stochastic-policy Actor-Critic:
Actor = policy network π_θ (updated with the clipped objective)
Critic = value network V_φ (provides the GAE advantage)
Relationship = standard AC architecture, with the Actor's update swapped to the clipped formHow to choose between PPO and SAC:
- Plenty of parallel compute + a stable deployment → PPO (on-policy stability, no off-policy distribution shift, easy deployment);
- Limited sample budget + expensive data collection → SAC (off-policy efficiency);
- Based on practical experience from tuning and hyperparameter optimization: beginners doing continuous control should start with SAC; discrete / multi-agent / large-scale parallel should start with PPO.
7. The Family Tree and a Unified View
text
┌─────────────┐
│ Policy │
│ Gradient │
│ REINFORCE │
└──────┬──────┘
│ + value baseline (Critic)
┌────────────────┼─────────────────┐
▼ ▼ ▼
┌───────────┐ ┌────────────┐ ┌───────────────┐
│ A2C/A3C │ │ Vanilla AC │ │ DPG │
│ parallel │ │ │ │ (deterministic│
│ envs │ │ │ │ policy grad) │
└───────────┘ └────────────┘ └───────┬───────┘
│ │ + DNN
▼ ▼
┌───────────┐ ┌───────────┐
│ PPO │ │ DDPG │
│ clipped │ │ continuous│
│ objective │ │ control │
└───────────┘ └─────┬─────┘
│ fix overest. /
│ delayed updates
▼
┌───────────┐
│ TD3 │
│ twin Q / │
│ delay / │
│ smoothing │
└─────┬─────┘
│ + entropy reg.
▼
┌───────────┐
│ SAC │
│ max-ent / │
│ auto α │
└───────────┘A unified view: every Actor-Critic answers two questions — "how good is the current action?" (the Critic, a value signal) and "what action should come next?" (the Actor, the policy). They differ in only three ways:
- Is the policy stochastic or deterministic? (Gaussian / categorical vs. direct action output);
- Is the data on-policy or off-policy? (use-once-and-discard vs. replay reuse);
- How is the update step constrained? (PPO uses clipping, TRPO uses KL, SAC uses entropy + soft updates).
8. Algorithm Selection and Engineering Notes
1. A Decision Tree for Choosing
text
Here's the problem
├── Discrete actions (games, menus) → DQN family / PPO
├── Continuous actions
│ ├── Cheap samples (unlimited simulation) → PPO (parallel rollouts)
│ ├── Expensive samples (real hardware, costly sim) → SAC
│ └── Need deterministic execution (no randomness at deploy) → TD3
└── Need an explicitly stochastic policy (games, exploration) → PPO / SAC2. Engineering Checklist
| Item | Advice |
|---|---|
| Let the Critic warm up | In off-policy algorithms the Critic is blind for the first few thousand steps — don't judge results too early |
| Reward scaling | SAC is sensitive to reward magnitude; normalize rewards first (see reward engineering) |
| Target network update rate τ | Around 0.005 by default; too large a τ makes the Critic unstable |
| Framework choice | Use a mature implementation instead of hand-rolling one: see How to Choose Frameworks and Tools |
| Multi-seed validation | AC algorithms are highly stochastic; single-seed conclusions are unreliable: see Building an RL Evaluation from Scratch |
The biggest pitfall — switching algorithms before the reward is designed
"SAC isn't working, switch to PPO; PPO isn't working, switch to TD3" is the most common futile loop. Most "the algorithm won't converge" cases are actually reward-design or environment-bug problems (see Common Pitfalls and Anti-Patterns). Run the ten-item checklist in reward engineering first, then talk about switching algorithms.
9. Where This Meets the Real World
- SAC and PPO are the stars of continuous control — see Robotics Control and Sim2Real (dexterous hands and the ANYmal quadruped both run on them);
- The foundational multi-agent algorithms MADDPG and MAPPO are both extensions of Actor-Critic — see multi-agent RL;
- The PPO phase of LLM alignment is Actor-Critic too (the policy is the LLM, the Critic is the reward model / value head) — see RLHF and Alignment with Human Feedback.
Further Reading
- Policy Gradient Methods — the theoretical foundation of the Actor side: the policy gradient theorem, GAE, PPO clipping
- Value-Based Learning: From Dynamic Programming to DQN — the foundation of the Critic side: TD error, experience replay, target networks
- Robotics Control and Sim2Real — engineering details of SAC/PPO on real robots
- How to Choose Frameworks and Tools — how to pick SAC and PPO implementations in SB3/Ray/Brax
- Tuning and Hyperparameter Optimization — the core hyperparameters of AC algorithms and the order in which to tune them
References
- Mnih, V., Badia, A. P., Mirza, M., et al. (2016). Asynchronous Methods for Deep Reinforcement Learning. ICML. arXiv:1602.01783 (A3C)
- Lillicrap, T. P., Hunt, J. J., Pritzel, A., et al. (2016). Continuous control with deep reinforcement learning. ICLR. arXiv:1509.02971 (DDPG)
- Fujimoto, S., van Hoof, H., & Meger, D. (2018). Addressing Function Approximation Error in Actor-Critic Methods. ICML. arXiv:1802.09477 (TD3)
- Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. ICML. arXiv:1801.01290 (SAC)
- Haarnoja, T., Zhou, A., Hartikainen, K., et al. (2018). Soft Actor-Critic Algorithms and Applications. arXiv:1812.05905 (SAC with automatic temperature)
- Konda, V. R., & Tsitsiklis, J. N. (2000). Actor-Critic Algorithms. NeurIPS. The theoretical foundation of Actor-Critic.