Appearance
Progressive Tutorial: Three Gymnasium Versions from Scratch
In one sentence: this page uses the same CartPole problem to walk you through writing three representative algorithms from scratch — tabular Q-learning, REINFORCE, and PPO — and getting each one to actually run. It bridges the "I understand the theory, but my code throws errors" gap, and it's aimed at readers who have just finished the concept pages and are about to write RL code for the first time. By the end, you'll have three working programs, a three-way algorithm comparison table, and a minimal reproducible experiment template.
All the code on this page is built on Gymnasium, the community-maintained successor to OpenAI Gym. The three versions share the same environment interface but learn in completely different ways:
CartPole evolution across three versions
┌──────────────────────────────────────────────────────────────┐
│ Version 1 Tabular Q-learning value-based: a lookup table, │
│ │ observations discretized │
│ Version 2 REINFORCE policy gradient: learn a │
│ │ network π(a|s) │
│ Version 3 Minimal PPO value + policy: Actor-Critic │
│ │ + GAE │
│ Common ground: same env API, same "reach 475" acceptance bar │
└──────────────────────────────────────────────────────────────┘One thing up front: this is not a three-way choice — it's three milestones. Write Q-learning with your own hands and you'll truly understand value overestimation and the price of discretization; write REINFORCE and you'll truly understand what "high variance" means; write PPO and you can honestly say you've crossed the threshold of modern deep RL. The theory behind the three versions lives in Value-Based Learning, Policy Gradient Methods, and The Actor-Critic Family — cross-reference them as you code.
1. Installation and a Gymnasium API Tour
1. Installation
bash
# Core environment
pip install gymnasium numpy
# PyTorch (needed for Versions 2 and 3; the CPU build is enough for CartPole)
pip install torch --index-url https://download.pytorch.org/whl/cpu # example: CPU build
# Optional: plot learning curves
pip install matplotlibPin your versions
The RL ecosystem is picky about versions. This page targets gymnasium >= 0.28 and numpy >= 1.24. After installing, run python -c "import gymnasium; print(gymnasium.__version__)" to confirm. The full index of environment libraries is in the Datasets & Tools Index.
2. A Minimal API Probe
First, get the "API probe" below running — it verifies, in one shot, the names and return structures of every API you'll touch:
python
import gymnasium as gym
# Create the environment; render_mode="human" opens a window
env = gym.make("CartPole-v1", render_mode="rgb_array")
# Observation space and action space
print("observation_space:", env.observation_space)
print("action_space:", env.action_space)
# reset: the new API returns a two-tuple (obs, info)
obs, info = env.reset(seed=42)
print("obs shape after reset:", obs.shape, "dtype:", obs.dtype)
print("info after reset:", info)
# Sample a random action and take one step
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
print(f"step returns: obs={obs}, reward={reward}")
print(f"terminated={terminated}, truncated={truncated}, info={info}")
env.close()The output tells you several important things:
| Fact | Value | Why it matters |
|---|---|---|
observation_space | Box([-4.8 -inf -0.418 -inf], [4.8 inf 0.418 inf], (4,), float32) | Continuous observations, 4-D, float32 |
action_space | Discrete(2) | Discrete actions: push left / push right |
reset | Returns a two-tuple (obs, info) | APIs differ across versions — check the return structure |
step | Returns a 5-tuple including terminated and truncated | Termination and truncation must be handled separately |
Two historical Gymnasium API pitfalls
- In older tutorials,
env.reset()returned justobs; since Gymnasium 0.26 it returns(obs, info). - Older tutorials used a 4-tuple
(obs, reward, done, info); Gymnasium splitdoneintoterminated(task completed) andtruncated(cut short by external forces, such as a time limit). In CartPole,truncatedis precisely the good outcome "survived all 500 steps" — treating it as failure destroys the learning signal. This is one of the most important pitfalls in this guide, and you'll see the correct handling in the code below.
3. About render
gym.make("CartPole-v1", render_mode="human") pops up a window for visualization; if you don't want a window, use "rgb_array" (which you can save as images or a GIF). Never use "human" during training — it slows things down dramatically. If you want to see the policy in action, run a few rendered evaluation episodes at the end of the training script.
2. The CartPole Problem
CartPole is the classic "first RL problem": a pole stands on a movable cart, and you push the cart left or right to keep the pole upright for as long as possible.
- State (observation): four real numbers — cart position
x, cart velocity, pole angleθfrom vertical, and angular velocity. - Actions: two discrete — push left / push right.
- Reward:
+1for every step you survive. - Termination: the pole tilts past ±12°, the cart leaves the track (±2.4), or you survive to step 500 (
truncated). - Goal: maximize cumulative reward; the theoretical ceiling is 500.
pole (angle θ closer to 0 is better)
│
┌────────┴───┐
│ cart │← force (left / right)
└────┬───────┘
└─────── track (episode ends if x leaves bounds)Why it makes a good tutorial: the state is only 4-D (small enough to tabulate), there are only 2 actions, an episode takes mere seconds, and you get results in minutes on a CPU. It's an extremely cheap probe for "is my training pipeline correct?" — any algorithm should comfortably reach 400+ on CartPole; if yours doesn't, the bug is in your implementation or your environment handling.
The shared acceptance bar for all three versions: a sliding average return of ≥ 475 over N consecutive episodes (i.e., near the 500 ceiling). If you don't hit it, debug your code first — don't blame the task for being "too hard."
3. Version 1: Tabular Q-Learning (Value-Based)
1. The Idea
Q-learning stores "state-action pair → value" in a table. But CartPole's observations are 4-D continuous values, so we first discretize: cut each dimension into bins, map a continuous observation to a bucket index, and the Q-table becomes a matrix of "bucket combination × action."
The update rule (off-policy TD control):
text
Q(s,a) ← Q(s,a) + α · [ r + γ · max_a' Q(s',a') − Q(s,a) ]The intuition: the TD error r + γ·max Q(s') − Q(s,a) says "the value after this step turned out better or worse than I thought," and you nudge the estimate a small step with learning rate α. For the full derivation, see Value-Based Learning: From Dynamic Programming to DQN.
2. Complete Code
python
"""
Version 1: tabular Q-learning on CartPole
Definition of done: sliding mean return >= 475 (near the 500 ceiling)
Run: python q_learning_cartpole.py
"""
import numpy as np
import gymnasium as gym
# ---------- 1. State discretization ----------
# Rough valid ranges for CartPole's 4 observation dims
OBS_BOUNDS = np.array([
[-4.8, 4.8], # cart position x
[-3.0, 3.0], # velocity (clipped)
[-0.5, 0.5], # pole angle θ (beyond this range the episode always ends)
[-3.0, 3.0], # angular velocity (clipped)
])
# Bins per dimension: [position, velocity, angle, angular velocity]
NUM_BINS = np.array([10, 10, 10, 10])
def discretize(obs):
"""Map a continuous observation to a discrete index (0 .. prod(NUM_BINS)-1)."""
bins = []
for i, (low, high) in enumerate(OBS_BOUNDS):
# Clip into bounds, then bin by width
clipped = np.clip(obs[i], low, high)
idx = int((clipped - low) / (high - low) * NUM_BINS[i])
idx = min(idx, NUM_BINS[i] - 1)
bins.append(idx)
# Encode as a single index (row-major order)
index = 0
for i, b in enumerate(bins):
index = index * NUM_BINS[i] + b
return index
# ---------- 2. Hyperparameters ----------
ALPHA = 0.1 # learning rate
GAMMA = 0.99 # discount factor
EPS_START = 1.0 # initial exploration rate
EPS_END = 0.01 # minimum exploration rate
EPS_DECAY = 0.9995 # decay per episode
NUM_EPISODES = 8000
MAX_STEPS = 500 # matches CartPole-v1's truncation limit
env = gym.make("CartPole-v1")
n_state = int(np.prod(NUM_BINS))
n_action = env.action_space.n
Q = np.zeros((n_state, n_action))
# ---------- 3. Training loop ----------
episode_returns = []
for ep in range(NUM_EPISODES):
obs, _ = env.reset(seed=ep) # seed every episode for reproducibility
state = discretize(obs)
eps = max(EPS_END, EPS_START * (EPS_DECAY ** ep))
total_reward = 0
for t in range(MAX_STEPS):
# ε-greedy action selection
if np.random.rand() < eps:
action = env.action_space.sample()
else:
action = int(np.argmax(Q[state]))
obs, reward, terminated, truncated, _ = env.step(action)
next_state = discretize(obs)
# Q-learning update (off-policy: uses the max next-state value directly)
best_next = np.max(Q[next_state])
Q[state, action] += ALPHA * (reward + GAMMA * best_next - Q[state, action])
state = next_state
total_reward += reward
if terminated or truncated:
break
episode_returns.append(total_reward)
if (ep + 1) % 500 == 0:
mean = float(np.mean(episode_returns[-100:]))
print(f"episode {ep+1:>5d} | mean_return(last_100) = {mean:.1f} | eps = {eps:.3f}")
env.close()
print("Q-learning training done.")3. Expected Learning Curve and the Definition of Done
Plot the returns across these 8,000 episodes and you'll see the classic staircase ascent:
text
Return
500 ┤ ┌─────────────────── plateau (near ceiling)
│ ┌─────┘
400 ┤ ┌───┘
│ ┌──┘
300 ┤ ┌──┘
│ ┌─┘
200 ┤ ┌──┘
│ ┌─┘
100 ┤ ┌──┘
│ ┌──┘
0 ┤─┘
└───────────────────────────────────────────────▶ episodes
0 1000 2000 3000 4000 5000 6000- First few hundred episodes: exploration dominates (ε is large) and returns hover around 10-50.
- From a few hundred to ~2,000 episodes: the Q-table starts earning its keep and returns climb in steps — note "steps," not a slope, because once the agent crosses a level, episode length suddenly jumps.
- After 2,000-3,000 episodes: returns approach the 500 ceiling and the sliding average settles above 475.
Want to see it fail once?
Set EPS_DECAY to 0.9999 (ε decays extremely slowly) and the first 2,000 episodes will stay stuck at low scores — a visceral demonstration of too much exploration suppressing performance, exactly the phenomenon described on the Exploration vs. Exploitation page.
Code-level pitfalls (specific to Version 1):
| Pitfall | Symptom | Fix |
|---|---|---|
discretize out of bounds | Some observations fall outside the bucket indices and Q indexing throws | Clip before computing the index (handled in the code) |
Treating truncated as terminated | Hitting 500 steps is misjudged as "failure," punishing learning in reverse | Break on both flags without distinguishing them — fine for this version |
| ε decays too fast | Turns greedy before exploring enough, stuck at low scores | Start at 1.0, decay exponentially to 0.01 |
Observation dtype is float32, compared directly in discretize | Precision issues shift bucket boundaries | Handle uniformly with np.clip (built in) |
4. Three Takeaways from Version 1
- Discretization is information loss: 10×10×10×10 = 10,000 cells, barely enough for CartPole; in high dimensions (like Atari frames) the cell count explodes exponentially — the embryo of the curse of dimensionality.
- Lookup tables don't generalize: a cell never visited is worth 0 forever, so you need dense discretization coverage.
- Value-based learning naturally fits discrete actions:
max_a Q(s,a)is cheap when there are few actions and painful when there are many — a hard constraint on the DQN family (Value-Based Learning).
4. Version 2: REINFORCE (Policy Gradient)
1. The Idea
Instead of a value table, we learn a policy network π_θ(a|s) directly: it takes the 4-D observation as input and outputs the probability of each of the 2 actions. Parameter updates follow the policy gradient theorem:
text
∇J(θ) = E_τ [ Σ_t G_t · ∇log π_θ(a_t|s_t) ]The intuition: whichever action led to a higher return G_t gets its probability pushed up. REINFORCE estimates G_t with the full-episode Monte Carlo return — no bootstrapping — so it's unbiased but high-variance. See Policy Gradient Methods for the derivation and intuition.
2. Complete Code
python
"""
Version 2: REINFORCE with discounted-return normalization (a baseline) on CartPole
Definition of done: sliding mean return >= 475
Run: python reinforce_cartpole.py
"""
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.distributions import Categorical
class PolicyNet(nn.Module):
"""Policy network: observation -> action probability distribution."""
def __init__(self, obs_dim, act_dim, hidden=128):
super().__init__()
self.net = nn.Sequential(
nn.Linear(obs_dim, hidden), nn.Tanh(),
nn.Linear(hidden, hidden), nn.Tanh(),
nn.Linear(hidden, act_dim), # outputs logits (no softmax)
)
def forward(self, obs):
return Categorical(logits=self.net(obs)) # applies softmax internally
def discount_returns(rewards, gamma=0.99):
"""Discounted returns G_t at each timestep (Monte Carlo, no bootstrapping)."""
returns = np.zeros(len(rewards), dtype=np.float32)
g = 0.0
for t in reversed(range(len(rewards))):
g = rewards[t] + gamma * g
returns[t] = g
return returns
def main(seed=0, num_episodes=3000, lr=1e-2, gamma=0.99):
torch.manual_seed(seed)
np.random.seed(seed)
env = gym.make("CartPole-v1")
policy = PolicyNet(env.observation_space.shape[0], env.action_space.n)
optimizer = optim.Adam(policy.parameters(), lr=lr)
episode_returns = []
for ep in range(num_episodes):
obs, _ = env.reset(seed=seed + ep)
log_probs, rewards = [], []
# 1) Sample a full episode
for _ in range(500):
obs_t = torch.as_tensor(obs, dtype=torch.float32)
dist = policy(obs_t)
action = dist.sample()
log_probs.append(dist.log_prob(action))
obs, reward, terminated, truncated, _ = env.step(action.item())
rewards.append(reward)
if terminated or truncated:
break
# 2) Compute returns and standardize (subtract mean, divide by std —
# the Monte Carlo version of a "baseline")
returns = torch.as_tensor(discount_returns(rewards, gamma), dtype=torch.float32)
returns = (returns - returns.mean()) / (returns.std() + 1e-8)
# 3) Policy gradient update (REINFORCE loss = - Σ G_t · log π(a_t|s_t))
loss = -torch.stack([
lp * ret for lp, ret in zip(log_probs, returns)
]).sum()
optimizer.zero_grad()
loss.backward()
optimizer.step()
episode_returns.append(sum(rewards))
if (ep + 1) % 100 == 0:
mean = float(np.mean(episode_returns[-100:]))
print(f"episode {ep+1:>5d} | mean_return(last_100) = {mean:.1f}")
env.close()
return episode_returns
if __name__ == "__main__":
main()3. Expected Learning Curve: The Variance King
REINFORCE's curve is violently jagged — that's its nature, not a bug:
text
Return
500 ┤ ┌──┐ ┌──┐
│ ┌──┘ └──┐ ┌──┘ └──┐
400 ┤──┘ └─┘ └────┐
│ └──┐ ┌────
300 ┤ └──┘
│
200 ┤
│
100 ┤
│
0 ┤
└──────────────────────────────────────────▶ episodes
0 1000 2000 3000- For the first few hundred episodes, returns oscillate wildly between 10 and 60.
- Around episodes 1,000-2,000, spikes suddenly touch 400 and then fall back — the direct signature of high gradient variance.
- By roughly episode 3,000, the sliding average (window 100) spends most of its time above 475.
High variance is what you're here to see
Same CartPole, same task, same goal — yet Q-learning climbs smoothly while REINFORCE climbs in convulsions. The only difference is the estimator. This is what "REINFORCE's Monte Carlo returns have high variance" looks like made flesh. Later, PPO presses the variance down with GAE and clipping, and you'll watch the curve visibly steady. To feel the variance even more deeply, raise lr from 1e-2 to 5e-2 and watch the curve go from "jittery" to "unhinged."
4. Pitfalls of Version 2
| Pitfall | Symptom | Fix |
|---|---|---|
| Returns not standardized | Gradient magnitudes unstable, extremely slow convergence | returns = (returns - mean)/(std+eps) |
Using mean() for loss instead of sum() | Gradient magnitude scales with episode length | Use sum, or normalize before mean() to keep magnitudes consistent |
| Learning rate too large (>1e-2) | One good episode gets overshot and rolled back; training collapses | Drop to 1e-2 or lower — the most sensitive knob in policy gradients |
torch.as_tensor(obs) without explicit dtype | Sporadic dtype-mismatch errors | Be explicit: dtype=torch.float32 |
Misremembering the return order of step | Unpacking errors | Memorize the 5-tuple (obs, reward, terminated, truncated, info) |
5. Three Takeaways from Version 2
- Continuous actions are no longer a problem: a policy network outputs distribution parameters directly, so REINFORCE handles continuous actions natively (with a Gaussian) — where tabular Q-learning is dead in the water.
- Variance is the Achilles' heel of policy gradients: the variance of full-episode returns grows linearly with episode length — the very reason TRPO and PPO exist (The Actor-Critic Family).
- The baseline idea: return standardization is the simplest baseline. A value network as an advantage baseline is the better answer — and that's Version 3.
5. Version 3: A Minimal PPO (Actor-Critic + GAE + Clipping)
1. The Idea
PPO adds three components on top of policy gradients, each with its own motivation:
| Component | What it does | Intuition |
|---|---|---|
| Actor-Critic | A value network supplies advantage estimates A(s,a) as a baseline | Replaces "how much absolute return did this step get" with "how much better than average was this step," reducing variance |
| GAE (generalized advantage estimation) | A tunable bias-variance trade-off | Large λ → closer to MC (high variance, low bias); small λ → closer to TD (low variance, high bias) |
| Clipping | Limits the step size of each update to prevent overshoot | A cheap approximation of a trust region that keeps training stable |
PPO's clipped surrogate objective:
text
L = min( r(θ)·A , clip(r(θ), 1−ε, 1+ε)·A )
where r(θ) = π_θ(a|s) / π_θ_old(a|s) (probability ratio of new to old policy)One sentence of intuition: only actions the old policy actually took are eligible for reinforcement, and each update round keeps the new-to-old probability ratio inside [1−ε, 1+ε]. For the formulas and mechanics, see Policy Gradient Methods and The Actor-Critic Family.
2. Complete Code (~150 lines, single environment)
python
"""
Version 3: minimal PPO on CartPole (single environment, no vectorization;
readability first)
Definition of done: sliding mean return >= 475
Run: python ppo_cartpole.py
"""
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.distributions import Categorical
class ActorCritic(nn.Module):
"""Shared feature layers + policy head + value head."""
def __init__(self, obs_dim, act_dim, hidden=128):
super().__init__()
self.shared = nn.Sequential(
nn.Linear(obs_dim, hidden), nn.Tanh(),
nn.Linear(hidden, hidden), nn.Tanh(),
)
self.policy_head = nn.Linear(hidden, act_dim) # logits
self.value_head = nn.Linear(hidden, 1) # scalar value V(s)
def dist(self, obs):
return Categorical(logits=self.policy_head(self.shared(obs)))
def value(self, obs):
return self.value_head(self.shared(obs)).squeeze(-1)
def compute_gae(rewards, values, next_value, dones, gamma, lam):
"""GAE: returns (advantages, returns). values does not include next_value."""
T = len(rewards)
advantages = torch.zeros(T)
last_gae = 0.0
for t in reversed(range(T)):
# if step t was truncated/terminated, no bootstrapping beyond it
mask = 1.0 - float(dones[t])
delta = rewards[t] + gamma * next_value * mask - values[t]
last_gae = delta + gamma * lam * mask * last_gae
advantages[t] = last_gae
next_value = values[t]
returns = advantages + values
return advantages, returns
def main(seed=0, total_steps=100_000, n_steps=256, gamma=0.99, lam=0.95,
lr=3e-4, clip_eps=0.2, update_epochs=10, minibatch_size=64,
ent_coef=0.01, vf_coef=0.5):
torch.manual_seed(seed)
np.random.seed(seed)
env = gym.make("CartPole-v1")
ac = ActorCritic(env.observation_space.shape[0], env.action_space.n)
optimizer = optim.Adam(ac.parameters(), lr=lr)
obs, _ = env.reset(seed=seed)
num_updates = total_steps // n_steps
for update in range(num_updates):
# ---------- 1) Sample a rollout ----------
# Collect Python scalars only; convert to tensors once at the end
# (clear, and avoids dtype traps)
obs_list, act_list, rew_list, dones_list, val_list = [], [], [], [], []
for _ in range(n_steps):
obs_t = torch.as_tensor(obs, dtype=torch.float32)
with torch.no_grad():
dist = ac.dist(obs_t)
value = ac.value(obs_t)
action = dist.sample()
obs_list.append(obs_t)
act_list.append(action)
val_list.append(value)
obs, reward, terminated, truncated, _ = env.step(action.item())
done = terminated or truncated
rew_list.append(float(reward))
dones_list.append(float(done))
if done:
obs, _ = env.reset()
# ---------- 2) GAE advantage computation ----------
obs_t = torch.as_tensor(obs, dtype=torch.float32)
with torch.no_grad():
next_value = ac.value(obs_t)
obs_tensor = torch.stack(obs_list)
with torch.no_grad():
old_dist = ac.dist(obs_tensor)
old_log_probs = old_dist.log_prob(torch.stack(act_list))
values_tensor = torch.stack(val_list)
rewards_tensor = torch.tensor(rew_list) # [T] one-shot conversion
dones_tensor = torch.tensor(dones_list) # [T]
advantages, returns = compute_gae(
rewards_tensor, values_tensor, next_value, dones_tensor, gamma, lam
)
advantages = (advantages - advantages.mean()) / (advantages.std() + 1e-8)
# ---------- 3) Multiple minibatch update epochs ----------
n_samples = n_steps
for _ in range(update_epochs):
perm = np.random.permutation(n_samples)
for start in range(0, n_samples, minibatch_size):
idx = perm[start:start + minibatch_size]
batch_obs = obs_tensor[idx]
batch_act = torch.stack(act_list)[idx]
batch_old = old_log_probs[idx]
batch_adv = advantages[idx]
batch_ret = returns[idx]
dist = ac.dist(batch_obs)
log_probs = dist.log_prob(batch_act)
entropy = dist.entropy().mean()
# probability ratio r(θ)
ratio = (log_probs - batch_old).exp()
# clipped surrogate objective
pg_loss = -torch.min(
ratio * batch_adv,
torch.clamp(ratio, 1.0 - clip_eps, 1.0 + clip_eps) * batch_adv,
).mean()
vf_loss = nn.functional.mse_loss(ac.value(batch_obs), batch_ret)
loss = pg_loss + vf_coef * vf_loss - ent_coef * entropy
optimizer.zero_grad()
loss.backward()
optimizer.step()
# ---------- 4) Logging ----------
total_reward = 0
eval_obs, _ = env.reset(seed=seed + 9999)
for _ in range(500):
with torch.no_grad():
eval_act = ac.dist(torch.as_tensor(eval_obs, dtype=torch.float32)).sample()
eval_obs, r, term, trunc, _ = env.step(eval_act.item())
total_reward += r
if term or trunc:
break
print(f"update {update+1:>4d}/{num_updates} | eval_return = {total_reward}")
env.close()
if __name__ == "__main__":
main()Two ways to collect tensors
The code above uses the clean approach: store Python scalars during collection and convert once with torch.tensor at the end. Another common style collects tensors directly and then torch.stacks them — both are fine, as long as dtypes stay consistent throughout. If you see torch.stack complaining about mismatched size or dtype while debugging, the usual culprit is a list mixing tensors of different shapes (e.g., a scalar reward stored alongside a (1,) tensor). Check your conversions line by line against the GAE section above.
A cleaner pattern, recommended for real use:
python
# Collect Python scalars during rollout
obs_list, act_list, rew_list, dones_list, val_list = [], [], [], [], []
for _ in range(n_steps):
...
rew_list.append(float(reward))
dones_list.append(float(done))
...
# Convert to tensors once after rollout
obs_tensor = torch.stack(obs_list) # [T, obs_dim]
act_tensor = torch.stack(act_list) # [T]
rew_tensor = torch.tensor(rew_list) # [T]
done_tensor = torch.tensor(dones_list) # [T]
val_tensor = torch.stack(val_list) # [T]3. Expected Learning Curve: Fast and Steady
text
Return
500 ┤ ┌─────────────── plateau
│ ┌─┘
400 ┤ ┌─┘
│ ┌──┘
300 ┤ ┌──┘
│ ┌──┘
200 ┤ ┌───┘
│ ┌──┘
100 ┤─┘
│
0 ┤
└──────────────────────────────────────────▶ updates
0 20 40 60 80 100 120Note the x-axis changed from "episodes" to "updates" (each update covers 256 steps). On CartPole, roughly 50-100 updates — about 13k-25k environment steps — reliably get you above 475: an order of magnitude better sample efficiency than REINFORCE, and the curve barely shows the kind of collapses REINFORCE suffers.
| Dimension | Q-learning | REINFORCE | Minimal PPO |
|---|---|---|---|
| Environment steps to reach 475 | ~20-30k | ~100-200k | ~15-30k |
| Curve stability | Smooth | Violently jagged | Steady |
| Lines of code | ~70 | ~80 | ~150 |
| Theoretical basis | Value-based | Policy gradient | Actor-Critic + GAE |
| Continuous actions? | No | Yes | Yes |
4. Pitfalls of Version 3 (Twice as Many as the First Two)
| Pitfall | Symptom | Fix |
|---|---|---|
old_log_probs not computed with the sampling-time policy | Probability ratios are distorted and training destabilizes | Must be computed with the frozen old policy before the update, under no_grad |
Forgetting the mask on GAE's dones | Bootstrapping wrongly crosses episode boundaries, polluting value estimates | delta = r + γ·V'·mask − V |
| Advantages not normalized | Advantage magnitudes vary across episodes, making the learning rate hard to tune | (A−mean)/(std+eps) |
Running a single environment continuously and forgetting reset | Observations reuse the terminal frame, corrupting value predictions | Call reset immediately after done |
| Entropy coefficient set to 0 | Acceptable on CartPole, but complex environments tend to collapse prematurely to a suboptimal policy | Keep a small entropy bonus: ent_coef=0.01 |
Action passed to env.step is a torch.Tensor | Gymnasium raises TypeError | Convert with .item() to a Python int |
Want to skip the handwriting and cross-check with SB3?
Pull out Stable-Baselines3 from the framework comparison and a single line — PPO("MlpPolicy", "CartPole-v1").learn(50000) — gives you the same result. The value of the hand-written version is that you can read every field in SB3's logs. Run the handwritten version first, then adopt the framework; walk on both legs.
6. Comparison and How to Choose
| Dimension | V1 Q-learning | V2 REINFORCE | V3 PPO |
|---|---|---|---|
| What's learned | Q-table | Policy network | Policy + value network |
| Bootstraps? | Yes (TD) | No (pure MC) | Yes (GAE) |
| Variance | Low | Very high | Low |
| Bias | Yes (discretization + bootstrapping) | None (unbiased) | Yes (bootstrapping, tunable) |
| Observation requirements | Must be discretized | Continuous, used as-is | Continuous, used as-is |
| Action space | Discrete | Any | Any |
| Sample efficiency | Medium | Low | High |
| Offline updates | Yes (off-policy) | No (on-policy) | Yes (minibatch on-policy) |
| Typical use | Teaching; small discrete problems | Teaching; understanding policy gradients | The default choice in modern RL |
How an engineer should choose:
- For learning: write all three, in the order of this page.
- For small discrete tasks (board games, simple scheduling): the Q-learning family is enough and extremely stable.
- For real business problems (continuous control, robotics, combinatorial optimization): go straight to PPO/SAC (The Actor-Critic Family).
- To validate a pipeline quickly: SB3's PPO produces a baseline in five minutes — then decide whether you need to build your own.
7. Master Pitfall Table (Applies to All Three Versions)
Every one of these reproduces on CartPole, but they usually only cost you dearly in real projects:
| # | Pitfall | Symptom | Detection / fix |
|---|---|---|---|
| 1 | Inconsistent observation dtypes (float32/float64 mixed) | Sporadic type errors or precision drift | Standardize on torch.float32 |
| 2 | Conflating terminated/truncated | Timeouts treated as failure, reversing the learning signal | Distinguish them wherever it matters (500 steps in CartPole is a happy ending) |
| 3 | Training with render_mode="human" | 10x slower training | Use rgb_array or no rendering in training; render only for evaluation |
| 4 | Seeding only one source of randomness | Environment and algorithm randomness blur together | Seed all three layers (Python/NumPy/torch + env); see Building an RL Project from Scratch |
| 5 | Judging from a single run's curve | Being fooled by variance | Run multiple seeds and plot IQR bands; see Building an RL Evaluation Suite from Scratch |
| 6 | Not calling env.close() | Leaked environment handles in multiprocessing | Use try/finally or with gym.make(...) as env |
8. A Minimal Reproducible Experiment Template
Once all three versions run, save the template below as run_experiment.py — it becomes the starting point for every experiment that follows, turning "train + evaluate + archive" into a single command:
python
"""Minimal reproducible experiment template: run PPO-CartPole once and save all artifacts."""
import argparse
import json
import numpy as np
import gymnasium as gym
import torch
from ppo_cartpole import ActorCritic # reuse the Version-3 model
def evaluate(ac, seed, episodes=10):
env = gym.make("CartPole-v1")
returns = []
for i in range(episodes):
obs, _ = env.reset(seed=seed * 100 + i)
total = 0
for _ in range(500):
with torch.no_grad():
act = ac.dist(torch.as_tensor(obs, dtype=torch.float32)).sample()
obs, r, term, trunc, _ = env.step(act.item())
total += r
if term or trunc:
break
returns.append(total)
env.close()
return np.mean(returns), np.std(returns)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--seed", type=int, default=0)
parser.add_argument("--lr", type=float, default=3e-4)
parser.add_argument("--total_steps", type=int, default=50_000)
args = parser.parse_args()
# Record the full config (the experiment's ID card)
config = vars(args)
torch.manual_seed(args.seed); np.random.seed(args.seed)
ac = ActorCritic(4, 2)
# ... plug in the Version-3 training loop here ...
mean, std = evaluate(ac, args.seed)
# Artifacts: model + config + metrics, all saved to disk
torch.save(ac.state_dict(), f"model_seed{args.seed}.pt")
with open(f"result_seed{args.seed}.json", "w", encoding="utf-8") as f:
json.dump({**config, "eval_mean": mean, "eval_std": std}, f, ensure_ascii=False, indent=2)
print(f"seed={args.seed} eval_return={mean:.1f} ± {std:.1f}")
if __name__ == "__main__":
main()Run three different seeds and merge the results:
bash
for s in 0 1 2; do python run_experiment.py --seed $s; doneThat's the minimal closed loop — multi-seed, reproducible, artifacts on disk. It is exactly the smallest block of the experiment matrix in Evaluation in Practice. On top of it, add Hydra config management and W&B/TensorBoard logging, and you have an engineering-grade experiment system.
9. Extending to Other Environments
Once CartPole works, swap the environment name in gym.make to try harder tasks:
| Environment | Action space | Difficulty jump | What to change |
|---|---|---|---|
LunarLander-v2 | Discrete, 4 | Sparser rewards, contact dynamics | Almost nothing (all three versions handle discrete actions) |
MountainCar-v0 | Discrete, 3 | Extremely sparse reward — plain Q-learning will struggle | Q-learning needs reward shaping or a curriculum |
Acrobot-v1 | Discrete, 3 | Continuous swing-up control | No code changes; tune hyperparameters |
Pendulum-v1 | Continuous, 1 | Your first continuous-action task | In Version 3, swap Categorical for a Gaussian |
HalfCheetah-v4/Humanoid-v4 (MuJoCo) | Continuous, 6/17 | High-dimensional continuous control | Needs vectorized sampling + GPU; a framework (SB3/Brax) is the pragmatic choice |
Every new environment gets the same acceptance routine
A new environment counts as "working" when: the random baseline scores → a simple policy learns something → the three-way comparison table reproduces the "PPO is fast and steady" pattern. This is the minimal version of the "build baselines before swapping algorithms" practice emphasized in Evaluation in Practice.
For more environment libraries (Brax, MuJoCo, Procgen, Atari, and more), see the Datasets & Tools Index; for the algorithm families behind the three versions, see Value-Based Learning, Policy Gradient Methods, and The Actor-Critic Family.
Further Reading
- Value-Based Learning: From Dynamic Programming to DQN — the full theory behind Version 1: from TD to the DQN family
- Policy Gradient Methods — the full theory behind Version 2: REINFORCE, baselines, the PPO objective
- The Actor-Critic Family — the full theory behind Version 3: GAE and the SAC/DDPG lineage
- Building an RL Project from Scratch — upgrade this page's "single experiment" into a "complete project pipeline"
- Building an RL Evaluation Suite from Scratch — the multi-seed, stricter version of this page's acceptance bar
- How to Choose Frameworks and Tools — after the handwritten versions run, how to adopt SB3 / RLlib / CleanRL
References
- Gymnasium official documentation (Farama Foundation): gymnasium.farama.org; the
Envinterface is documented at gymnasium.farama.org/api/env/ - Watkins, C. J. C. H. & Dayan, P. (1992). Q-Learning. Machine Learning, 8, 279–292. link.springer.com/article/10.1007/BF00992698
- Williams, R. J. (1992). Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning, 8, 229–256. link.springer.com/article/10.1007/BF00992696
- Schulman, J. et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347
- Schulman, J. et al. (2016). High-Dimensional Continuous Control Using Generalized Advantage Estimation. ICLR 2016. arXiv:1506.02438
- CleanRL official repository and documentation: github.com/vwxyzjn/cleanrl, docs.cleanrl.dev (this page's PPO structure follows CleanRL's minimalist style)
- Stable-Baselines3 official documentation: stable-baselines3.readthedocs.io (a mature reference implementation of PPO)