Skip to content

Progressive Tutorial: Three Gymnasium Versions from Scratch

On this page Implement three versions of CartPole from scratch with Gymnasium — tabular Q-learning, a simple policy gradient, then PPO. Each version ships with complete code, key pitfalls, and a "definition of done," plus a minimal reproducible experiment template.

Progressive Tutorial: Three Gymnasium Versions from Scratch ​

In one sentence: this page uses the same CartPole problem to walk you through writing three representative algorithms from scratch — tabular Q-learning, REINFORCE, and PPO — and getting each one to actually run. It bridges the "I understand the theory, but my code throws errors" gap, and it's aimed at readers who have just finished the concept pages and are about to write RL code for the first time. By the end, you'll have three working programs, a three-way algorithm comparison table, and a minimal reproducible experiment template.

All the code on this page is built on Gymnasium, the community-maintained successor to OpenAI Gym. The three versions share the same environment interface but learn in completely different ways:

CartPole evolution across three versions
┌──────────────────────────────────────────────────────────────┐
│ Version 1  Tabular Q-learning   value-based: a lookup table, │
│    │                              observations discretized    │
│ Version 2  REINFORCE            policy gradient: learn a     │
│    │                              network π(a|s)              │
│ Version 3  Minimal PPO          value + policy: Actor-Critic │
│    │                              + GAE                      │
│ Common ground: same env API, same "reach 475" acceptance bar │
└──────────────────────────────────────────────────────────────┘

One thing up front: this is not a three-way choice — it's three milestones. Write Q-learning with your own hands and you'll truly understand value overestimation and the price of discretization; write REINFORCE and you'll truly understand what "high variance" means; write PPO and you can honestly say you've crossed the threshold of modern deep RL. The theory behind the three versions lives in Value-Based Learning, Policy Gradient Methods, and The Actor-Critic Family — cross-reference them as you code.

1. Installation and a Gymnasium API Tour ​

1. Installation ​

bash
# Core environment
pip install gymnasium numpy
# PyTorch (needed for Versions 2 and 3; the CPU build is enough for CartPole)
pip install torch --index-url https://download.pytorch.org/whl/cpu   # example: CPU build
# Optional: plot learning curves
pip install matplotlib

Pin your versions

The RL ecosystem is picky about versions. This page targets gymnasium >= 0.28 and numpy >= 1.24. After installing, run python -c "import gymnasium; print(gymnasium.__version__)" to confirm. The full index of environment libraries is in the Datasets & Tools Index.

2. A Minimal API Probe ​

First, get the "API probe" below running — it verifies, in one shot, the names and return structures of every API you'll touch:

python
import gymnasium as gym

# Create the environment; render_mode="human" opens a window
env = gym.make("CartPole-v1", render_mode="rgb_array")

# Observation space and action space
print("observation_space:", env.observation_space)
print("action_space:", env.action_space)

# reset: the new API returns a two-tuple (obs, info)
obs, info = env.reset(seed=42)
print("obs shape after reset:", obs.shape, "dtype:", obs.dtype)
print("info after reset:", info)

# Sample a random action and take one step
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
print(f"step returns: obs={obs}, reward={reward}")
print(f"terminated={terminated}, truncated={truncated}, info={info}")

env.close()

The output tells you several important things:

FactValueWhy it matters
observation_spaceBox([-4.8 -inf -0.418 -inf], [4.8 inf 0.418 inf], (4,), float32)Continuous observations, 4-D, float32
action_spaceDiscrete(2)Discrete actions: push left / push right
resetReturns a two-tuple (obs, info)APIs differ across versions — check the return structure
stepReturns a 5-tuple including terminated and truncatedTermination and truncation must be handled separately

Two historical Gymnasium API pitfalls

  1. In older tutorials, env.reset() returned just obs; since Gymnasium 0.26 it returns (obs, info).
  2. Older tutorials used a 4-tuple (obs, reward, done, info); Gymnasium split done into terminated (task completed) and truncated (cut short by external forces, such as a time limit). In CartPole, truncated is precisely the good outcome "survived all 500 steps" — treating it as failure destroys the learning signal. This is one of the most important pitfalls in this guide, and you'll see the correct handling in the code below.

3. About render ​

gym.make("CartPole-v1", render_mode="human") pops up a window for visualization; if you don't want a window, use "rgb_array" (which you can save as images or a GIF). Never use "human" during training — it slows things down dramatically. If you want to see the policy in action, run a few rendered evaluation episodes at the end of the training script.

2. The CartPole Problem ​

CartPole is the classic "first RL problem": a pole stands on a movable cart, and you push the cart left or right to keep the pole upright for as long as possible.

  • State (observation): four real numbers — cart position x, cart velocity, pole angle θ from vertical, and angular velocity.
  • Actions: two discrete — push left / push right.
  • Reward: +1 for every step you survive.
  • Termination: the pole tilts past ±12°, the cart leaves the track (±2.4), or you survive to step 500 (truncated).
  • Goal: maximize cumulative reward; the theoretical ceiling is 500.
             pole (angle θ closer to 0 is better)
              │
     ┌────────┴───┐
     │   cart     │← force (left / right)
     └────┬───────┘
          └─────── track (episode ends if x leaves bounds)

Why it makes a good tutorial: the state is only 4-D (small enough to tabulate), there are only 2 actions, an episode takes mere seconds, and you get results in minutes on a CPU. It's an extremely cheap probe for "is my training pipeline correct?" — any algorithm should comfortably reach 400+ on CartPole; if yours doesn't, the bug is in your implementation or your environment handling.

The shared acceptance bar for all three versions: a sliding average return of ≥ 475 over N consecutive episodes (i.e., near the 500 ceiling). If you don't hit it, debug your code first — don't blame the task for being "too hard."

3. Version 1: Tabular Q-Learning (Value-Based) ​

1. The Idea ​

Q-learning stores "state-action pair → value" in a table. But CartPole's observations are 4-D continuous values, so we first discretize: cut each dimension into bins, map a continuous observation to a bucket index, and the Q-table becomes a matrix of "bucket combination × action."

The update rule (off-policy TD control):

text
Q(s,a) ← Q(s,a) + α · [ r + γ · max_a' Q(s',a') − Q(s,a) ]

The intuition: the TD error r + γ·max Q(s') − Q(s,a) says "the value after this step turned out better or worse than I thought," and you nudge the estimate a small step with learning rate α. For the full derivation, see Value-Based Learning: From Dynamic Programming to DQN.

2. Complete Code ​

python
"""
Version 1: tabular Q-learning on CartPole
Definition of done: sliding mean return >= 475 (near the 500 ceiling)
Run: python q_learning_cartpole.py
"""
import numpy as np
import gymnasium as gym

# ---------- 1. State discretization ----------
# Rough valid ranges for CartPole's 4 observation dims
OBS_BOUNDS = np.array([
    [-4.8, 4.8],      # cart position x
    [-3.0, 3.0],      # velocity (clipped)
    [-0.5, 0.5],      # pole angle θ (beyond this range the episode always ends)
    [-3.0, 3.0],      # angular velocity (clipped)
])
# Bins per dimension: [position, velocity, angle, angular velocity]
NUM_BINS = np.array([10, 10, 10, 10])


def discretize(obs):
    """Map a continuous observation to a discrete index (0 .. prod(NUM_BINS)-1)."""
    bins = []
    for i, (low, high) in enumerate(OBS_BOUNDS):
        # Clip into bounds, then bin by width
        clipped = np.clip(obs[i], low, high)
        idx = int((clipped - low) / (high - low) * NUM_BINS[i])
        idx = min(idx, NUM_BINS[i] - 1)
        bins.append(idx)
    # Encode as a single index (row-major order)
    index = 0
    for i, b in enumerate(bins):
        index = index * NUM_BINS[i] + b
    return index


# ---------- 2. Hyperparameters ----------
ALPHA = 0.1            # learning rate
GAMMA = 0.99           # discount factor
EPS_START = 1.0        # initial exploration rate
EPS_END = 0.01         # minimum exploration rate
EPS_DECAY = 0.9995     # decay per episode
NUM_EPISODES = 8000
MAX_STEPS = 500        # matches CartPole-v1's truncation limit

env = gym.make("CartPole-v1")
n_state = int(np.prod(NUM_BINS))
n_action = env.action_space.n
Q = np.zeros((n_state, n_action))

# ---------- 3. Training loop ----------
episode_returns = []
for ep in range(NUM_EPISODES):
    obs, _ = env.reset(seed=ep)          # seed every episode for reproducibility
    state = discretize(obs)
    eps = max(EPS_END, EPS_START * (EPS_DECAY ** ep))
    total_reward = 0

    for t in range(MAX_STEPS):
        # ε-greedy action selection
        if np.random.rand() < eps:
            action = env.action_space.sample()
        else:
            action = int(np.argmax(Q[state]))

        obs, reward, terminated, truncated, _ = env.step(action)
        next_state = discretize(obs)

        # Q-learning update (off-policy: uses the max next-state value directly)
        best_next = np.max(Q[next_state])
        Q[state, action] += ALPHA * (reward + GAMMA * best_next - Q[state, action])

        state = next_state
        total_reward += reward
        if terminated or truncated:
            break

    episode_returns.append(total_reward)

    if (ep + 1) % 500 == 0:
        mean = float(np.mean(episode_returns[-100:]))
        print(f"episode {ep+1:>5d} | mean_return(last_100) = {mean:.1f} | eps = {eps:.3f}")

env.close()
print("Q-learning training done.")

3. Expected Learning Curve and the Definition of Done ​

Plot the returns across these 8,000 episodes and you'll see the classic staircase ascent:

text
Return
500 ┤                              ┌─────────────────── plateau (near ceiling)
    │                        ┌─────┘
400 ┤                    ┌───┘
    │                 ┌──┘
300 ┤              ┌──┘
    │            ┌─┘
200 ┤         ┌──┘
    │       ┌─┘
100 ┤    ┌──┘
    │ ┌──┘
  0 ┤─┘
    └───────────────────────────────────────────────▶ episodes
     0    1000   2000   3000   4000   5000   6000
  • First few hundred episodes: exploration dominates (ε is large) and returns hover around 10-50.
  • From a few hundred to ~2,000 episodes: the Q-table starts earning its keep and returns climb in steps — note "steps," not a slope, because once the agent crosses a level, episode length suddenly jumps.
  • After 2,000-3,000 episodes: returns approach the 500 ceiling and the sliding average settles above 475.

Want to see it fail once?

Set EPS_DECAY to 0.9999 (ε decays extremely slowly) and the first 2,000 episodes will stay stuck at low scores — a visceral demonstration of too much exploration suppressing performance, exactly the phenomenon described on the Exploration vs. Exploitation page.

Code-level pitfalls (specific to Version 1):

PitfallSymptomFix
discretize out of boundsSome observations fall outside the bucket indices and Q indexing throwsClip before computing the index (handled in the code)
Treating truncated as terminatedHitting 500 steps is misjudged as "failure," punishing learning in reverseBreak on both flags without distinguishing them — fine for this version
ε decays too fastTurns greedy before exploring enough, stuck at low scoresStart at 1.0, decay exponentially to 0.01
Observation dtype is float32, compared directly in discretizePrecision issues shift bucket boundariesHandle uniformly with np.clip (built in)

4. Three Takeaways from Version 1 ​

  1. Discretization is information loss: 10×10×10×10 = 10,000 cells, barely enough for CartPole; in high dimensions (like Atari frames) the cell count explodes exponentially — the embryo of the curse of dimensionality.
  2. Lookup tables don't generalize: a cell never visited is worth 0 forever, so you need dense discretization coverage.
  3. Value-based learning naturally fits discrete actions: max_a Q(s,a) is cheap when there are few actions and painful when there are many — a hard constraint on the DQN family (Value-Based Learning).

4. Version 2: REINFORCE (Policy Gradient) ​

1. The Idea ​

Instead of a value table, we learn a policy network π_θ(a|s) directly: it takes the 4-D observation as input and outputs the probability of each of the 2 actions. Parameter updates follow the policy gradient theorem:

text
∇J(θ) = E_τ [ Σ_t G_t · ∇log π_θ(a_t|s_t) ]

The intuition: whichever action led to a higher return G_t gets its probability pushed up. REINFORCE estimates G_t with the full-episode Monte Carlo return — no bootstrapping — so it's unbiased but high-variance. See Policy Gradient Methods for the derivation and intuition.

2. Complete Code ​

python
"""
Version 2: REINFORCE with discounted-return normalization (a baseline) on CartPole
Definition of done: sliding mean return >= 475
Run: python reinforce_cartpole.py
"""
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.distributions import Categorical


class PolicyNet(nn.Module):
    """Policy network: observation -> action probability distribution."""
    def __init__(self, obs_dim, act_dim, hidden=128):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(obs_dim, hidden), nn.Tanh(),
            nn.Linear(hidden, hidden), nn.Tanh(),
            nn.Linear(hidden, act_dim),   # outputs logits (no softmax)
        )

    def forward(self, obs):
        return Categorical(logits=self.net(obs))   # applies softmax internally


def discount_returns(rewards, gamma=0.99):
    """Discounted returns G_t at each timestep (Monte Carlo, no bootstrapping)."""
    returns = np.zeros(len(rewards), dtype=np.float32)
    g = 0.0
    for t in reversed(range(len(rewards))):
        g = rewards[t] + gamma * g
        returns[t] = g
    return returns


def main(seed=0, num_episodes=3000, lr=1e-2, gamma=0.99):
    torch.manual_seed(seed)
    np.random.seed(seed)

    env = gym.make("CartPole-v1")
    policy = PolicyNet(env.observation_space.shape[0], env.action_space.n)
    optimizer = optim.Adam(policy.parameters(), lr=lr)

    episode_returns = []
    for ep in range(num_episodes):
        obs, _ = env.reset(seed=seed + ep)
        log_probs, rewards = [], []

        # 1) Sample a full episode
        for _ in range(500):
            obs_t = torch.as_tensor(obs, dtype=torch.float32)
            dist = policy(obs_t)
            action = dist.sample()
            log_probs.append(dist.log_prob(action))
            obs, reward, terminated, truncated, _ = env.step(action.item())
            rewards.append(reward)
            if terminated or truncated:
                break

        # 2) Compute returns and standardize (subtract mean, divide by std —
        #    the Monte Carlo version of a "baseline")
        returns = torch.as_tensor(discount_returns(rewards, gamma), dtype=torch.float32)
        returns = (returns - returns.mean()) / (returns.std() + 1e-8)

        # 3) Policy gradient update (REINFORCE loss = - Σ G_t · log π(a_t|s_t))
        loss = -torch.stack([
            lp * ret for lp, ret in zip(log_probs, returns)
        ]).sum()

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        episode_returns.append(sum(rewards))
        if (ep + 1) % 100 == 0:
            mean = float(np.mean(episode_returns[-100:]))
            print(f"episode {ep+1:>5d} | mean_return(last_100) = {mean:.1f}")

    env.close()
    return episode_returns


if __name__ == "__main__":
    main()

3. Expected Learning Curve: The Variance King ​

REINFORCE's curve is violently jagged — that's its nature, not a bug:

text
Return
500 ┤     ┌──┐        ┌──┐
    │  ┌──┘  └──┐ ┌──┘  └──┐
400 ┤──┘        └─┘        └────┐
    │                            └──┐  ┌────
300 ┤                               └──┘
    │
200 ┤
    │
100 ┤
    │
  0 ┤
    └──────────────────────────────────────────▶ episodes
     0        1000        2000        3000
  • For the first few hundred episodes, returns oscillate wildly between 10 and 60.
  • Around episodes 1,000-2,000, spikes suddenly touch 400 and then fall back — the direct signature of high gradient variance.
  • By roughly episode 3,000, the sliding average (window 100) spends most of its time above 475.

High variance is what you're here to see

Same CartPole, same task, same goal — yet Q-learning climbs smoothly while REINFORCE climbs in convulsions. The only difference is the estimator. This is what "REINFORCE's Monte Carlo returns have high variance" looks like made flesh. Later, PPO presses the variance down with GAE and clipping, and you'll watch the curve visibly steady. To feel the variance even more deeply, raise lr from 1e-2 to 5e-2 and watch the curve go from "jittery" to "unhinged."

4. Pitfalls of Version 2 ​

PitfallSymptomFix
Returns not standardizedGradient magnitudes unstable, extremely slow convergencereturns = (returns - mean)/(std+eps)
Using mean() for loss instead of sum()Gradient magnitude scales with episode lengthUse sum, or normalize before mean() to keep magnitudes consistent
Learning rate too large (>1e-2)One good episode gets overshot and rolled back; training collapsesDrop to 1e-2 or lower — the most sensitive knob in policy gradients
torch.as_tensor(obs) without explicit dtypeSporadic dtype-mismatch errorsBe explicit: dtype=torch.float32
Misremembering the return order of stepUnpacking errorsMemorize the 5-tuple (obs, reward, terminated, truncated, info)

5. Three Takeaways from Version 2 ​

  1. Continuous actions are no longer a problem: a policy network outputs distribution parameters directly, so REINFORCE handles continuous actions natively (with a Gaussian) — where tabular Q-learning is dead in the water.
  2. Variance is the Achilles' heel of policy gradients: the variance of full-episode returns grows linearly with episode length — the very reason TRPO and PPO exist (The Actor-Critic Family).
  3. The baseline idea: return standardization is the simplest baseline. A value network as an advantage baseline is the better answer — and that's Version 3.

5. Version 3: A Minimal PPO (Actor-Critic + GAE + Clipping) ​

1. The Idea ​

PPO adds three components on top of policy gradients, each with its own motivation:

ComponentWhat it doesIntuition
Actor-CriticA value network supplies advantage estimates A(s,a) as a baselineReplaces "how much absolute return did this step get" with "how much better than average was this step," reducing variance
GAE (generalized advantage estimation)A tunable bias-variance trade-offLarge λ → closer to MC (high variance, low bias); small λ → closer to TD (low variance, high bias)
ClippingLimits the step size of each update to prevent overshootA cheap approximation of a trust region that keeps training stable

PPO's clipped surrogate objective:

text
L = min( r(θ)·A , clip(r(θ), 1−ε, 1+ε)·A )
where r(θ) = π_θ(a|s) / π_θ_old(a|s)  (probability ratio of new to old policy)

One sentence of intuition: only actions the old policy actually took are eligible for reinforcement, and each update round keeps the new-to-old probability ratio inside [1−ε, 1+ε]. For the formulas and mechanics, see Policy Gradient Methods and The Actor-Critic Family.

2. Complete Code (~150 lines, single environment) ​

python
"""
Version 3: minimal PPO on CartPole (single environment, no vectorization;
readability first)
Definition of done: sliding mean return >= 475
Run: python ppo_cartpole.py
"""
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.distributions import Categorical


class ActorCritic(nn.Module):
    """Shared feature layers + policy head + value head."""
    def __init__(self, obs_dim, act_dim, hidden=128):
        super().__init__()
        self.shared = nn.Sequential(
            nn.Linear(obs_dim, hidden), nn.Tanh(),
            nn.Linear(hidden, hidden), nn.Tanh(),
        )
        self.policy_head = nn.Linear(hidden, act_dim)  # logits
        self.value_head = nn.Linear(hidden, 1)         # scalar value V(s)

    def dist(self, obs):
        return Categorical(logits=self.policy_head(self.shared(obs)))

    def value(self, obs):
        return self.value_head(self.shared(obs)).squeeze(-1)


def compute_gae(rewards, values, next_value, dones, gamma, lam):
    """GAE: returns (advantages, returns). values does not include next_value."""
    T = len(rewards)
    advantages = torch.zeros(T)
    last_gae = 0.0
    for t in reversed(range(T)):
        # if step t was truncated/terminated, no bootstrapping beyond it
        mask = 1.0 - float(dones[t])
        delta = rewards[t] + gamma * next_value * mask - values[t]
        last_gae = delta + gamma * lam * mask * last_gae
        advantages[t] = last_gae
        next_value = values[t]
    returns = advantages + values
    return advantages, returns


def main(seed=0, total_steps=100_000, n_steps=256, gamma=0.99, lam=0.95,
         lr=3e-4, clip_eps=0.2, update_epochs=10, minibatch_size=64,
         ent_coef=0.01, vf_coef=0.5):
    torch.manual_seed(seed)
    np.random.seed(seed)

    env = gym.make("CartPole-v1")
    ac = ActorCritic(env.observation_space.shape[0], env.action_space.n)
    optimizer = optim.Adam(ac.parameters(), lr=lr)

    obs, _ = env.reset(seed=seed)
    num_updates = total_steps // n_steps

    for update in range(num_updates):
        # ---------- 1) Sample a rollout ----------
        # Collect Python scalars only; convert to tensors once at the end
        # (clear, and avoids dtype traps)
        obs_list, act_list, rew_list, dones_list, val_list = [], [], [], [], []
        for _ in range(n_steps):
            obs_t = torch.as_tensor(obs, dtype=torch.float32)
            with torch.no_grad():
                dist = ac.dist(obs_t)
                value = ac.value(obs_t)
                action = dist.sample()
            obs_list.append(obs_t)
            act_list.append(action)
            val_list.append(value)
            obs, reward, terminated, truncated, _ = env.step(action.item())
            done = terminated or truncated
            rew_list.append(float(reward))
            dones_list.append(float(done))
            if done:
                obs, _ = env.reset()

        # ---------- 2) GAE advantage computation ----------
        obs_t = torch.as_tensor(obs, dtype=torch.float32)
        with torch.no_grad():
            next_value = ac.value(obs_t)
        obs_tensor = torch.stack(obs_list)
        with torch.no_grad():
            old_dist = ac.dist(obs_tensor)
            old_log_probs = old_dist.log_prob(torch.stack(act_list))
            values_tensor = torch.stack(val_list)
        rewards_tensor = torch.tensor(rew_list)      # [T] one-shot conversion
        dones_tensor = torch.tensor(dones_list)      # [T]
        advantages, returns = compute_gae(
            rewards_tensor, values_tensor, next_value, dones_tensor, gamma, lam
        )
        advantages = (advantages - advantages.mean()) / (advantages.std() + 1e-8)

        # ---------- 3) Multiple minibatch update epochs ----------
        n_samples = n_steps
        for _ in range(update_epochs):
            perm = np.random.permutation(n_samples)
            for start in range(0, n_samples, minibatch_size):
                idx = perm[start:start + minibatch_size]
                batch_obs = obs_tensor[idx]
                batch_act = torch.stack(act_list)[idx]
                batch_old = old_log_probs[idx]
                batch_adv = advantages[idx]
                batch_ret = returns[idx]

                dist = ac.dist(batch_obs)
                log_probs = dist.log_prob(batch_act)
                entropy = dist.entropy().mean()

                # probability ratio r(θ)
                ratio = (log_probs - batch_old).exp()
                # clipped surrogate objective
                pg_loss = -torch.min(
                    ratio * batch_adv,
                    torch.clamp(ratio, 1.0 - clip_eps, 1.0 + clip_eps) * batch_adv,
                ).mean()
                vf_loss = nn.functional.mse_loss(ac.value(batch_obs), batch_ret)
                loss = pg_loss + vf_coef * vf_loss - ent_coef * entropy

                optimizer.zero_grad()
                loss.backward()
                optimizer.step()

        # ---------- 4) Logging ----------
        total_reward = 0
        eval_obs, _ = env.reset(seed=seed + 9999)
        for _ in range(500):
            with torch.no_grad():
                eval_act = ac.dist(torch.as_tensor(eval_obs, dtype=torch.float32)).sample()
            eval_obs, r, term, trunc, _ = env.step(eval_act.item())
            total_reward += r
            if term or trunc:
                break
        print(f"update {update+1:>4d}/{num_updates} | eval_return = {total_reward}")

    env.close()


if __name__ == "__main__":
    main()

Two ways to collect tensors

The code above uses the clean approach: store Python scalars during collection and convert once with torch.tensor at the end. Another common style collects tensors directly and then torch.stacks them — both are fine, as long as dtypes stay consistent throughout. If you see torch.stack complaining about mismatched size or dtype while debugging, the usual culprit is a list mixing tensors of different shapes (e.g., a scalar reward stored alongside a (1,) tensor). Check your conversions line by line against the GAE section above.

A cleaner pattern, recommended for real use:

python
# Collect Python scalars during rollout
obs_list, act_list, rew_list, dones_list, val_list = [], [], [], [], []
for _ in range(n_steps):
    ...
    rew_list.append(float(reward))
    dones_list.append(float(done))
...
# Convert to tensors once after rollout
obs_tensor  = torch.stack(obs_list)                 # [T, obs_dim]
act_tensor  = torch.stack(act_list)                 # [T]
rew_tensor  = torch.tensor(rew_list)                # [T]
done_tensor = torch.tensor(dones_list)              # [T]
val_tensor  = torch.stack(val_list)                 # [T]

3. Expected Learning Curve: Fast and Steady ​

text
Return
500 ┤                    ┌─────────────── plateau
    │                  ┌─┘
400 ┤                ┌─┘
    │             ┌──┘
300 ┤          ┌──┘
    │       ┌──┘
200 ┤    ┌───┘
    │ ┌──┘
100 ┤─┘
    │
  0 ┤
    └──────────────────────────────────────────▶ updates
     0     20     40     60     80    100    120

Note the x-axis changed from "episodes" to "updates" (each update covers 256 steps). On CartPole, roughly 50-100 updates — about 13k-25k environment steps — reliably get you above 475: an order of magnitude better sample efficiency than REINFORCE, and the curve barely shows the kind of collapses REINFORCE suffers.

DimensionQ-learningREINFORCEMinimal PPO
Environment steps to reach 475~20-30k~100-200k~15-30k
Curve stabilitySmoothViolently jaggedSteady
Lines of code~70~80~150
Theoretical basisValue-basedPolicy gradientActor-Critic + GAE
Continuous actions?NoYesYes

4. Pitfalls of Version 3 (Twice as Many as the First Two) ​

PitfallSymptomFix
old_log_probs not computed with the sampling-time policyProbability ratios are distorted and training destabilizesMust be computed with the frozen old policy before the update, under no_grad
Forgetting the mask on GAE's donesBootstrapping wrongly crosses episode boundaries, polluting value estimatesdelta = r + γ·V'·mask − V
Advantages not normalizedAdvantage magnitudes vary across episodes, making the learning rate hard to tune(A−mean)/(std+eps)
Running a single environment continuously and forgetting resetObservations reuse the terminal frame, corrupting value predictionsCall reset immediately after done
Entropy coefficient set to 0Acceptable on CartPole, but complex environments tend to collapse prematurely to a suboptimal policyKeep a small entropy bonus: ent_coef=0.01
Action passed to env.step is a torch.TensorGymnasium raises TypeErrorConvert with .item() to a Python int

Want to skip the handwriting and cross-check with SB3?

Pull out Stable-Baselines3 from the framework comparison and a single line — PPO("MlpPolicy", "CartPole-v1").learn(50000) — gives you the same result. The value of the hand-written version is that you can read every field in SB3's logs. Run the handwritten version first, then adopt the framework; walk on both legs.

6. Comparison and How to Choose ​

DimensionV1 Q-learningV2 REINFORCEV3 PPO
What's learnedQ-tablePolicy networkPolicy + value network
Bootstraps?Yes (TD)No (pure MC)Yes (GAE)
VarianceLowVery highLow
BiasYes (discretization + bootstrapping)None (unbiased)Yes (bootstrapping, tunable)
Observation requirementsMust be discretizedContinuous, used as-isContinuous, used as-is
Action spaceDiscreteAnyAny
Sample efficiencyMediumLowHigh
Offline updatesYes (off-policy)No (on-policy)Yes (minibatch on-policy)
Typical useTeaching; small discrete problemsTeaching; understanding policy gradientsThe default choice in modern RL

How an engineer should choose:

  1. For learning: write all three, in the order of this page.
  2. For small discrete tasks (board games, simple scheduling): the Q-learning family is enough and extremely stable.
  3. For real business problems (continuous control, robotics, combinatorial optimization): go straight to PPO/SAC (The Actor-Critic Family).
  4. To validate a pipeline quickly: SB3's PPO produces a baseline in five minutes — then decide whether you need to build your own.

7. Master Pitfall Table (Applies to All Three Versions) ​

Every one of these reproduces on CartPole, but they usually only cost you dearly in real projects:

#PitfallSymptomDetection / fix
1Inconsistent observation dtypes (float32/float64 mixed)Sporadic type errors or precision driftStandardize on torch.float32
2Conflating terminated/truncatedTimeouts treated as failure, reversing the learning signalDistinguish them wherever it matters (500 steps in CartPole is a happy ending)
3Training with render_mode="human"10x slower trainingUse rgb_array or no rendering in training; render only for evaluation
4Seeding only one source of randomnessEnvironment and algorithm randomness blur togetherSeed all three layers (Python/NumPy/torch + env); see Building an RL Project from Scratch
5Judging from a single run's curveBeing fooled by varianceRun multiple seeds and plot IQR bands; see Building an RL Evaluation Suite from Scratch
6Not calling env.close()Leaked environment handles in multiprocessingUse try/finally or with gym.make(...) as env

8. A Minimal Reproducible Experiment Template ​

Once all three versions run, save the template below as run_experiment.py — it becomes the starting point for every experiment that follows, turning "train + evaluate + archive" into a single command:

python
"""Minimal reproducible experiment template: run PPO-CartPole once and save all artifacts."""
import argparse
import json
import numpy as np
import gymnasium as gym
import torch
from ppo_cartpole import ActorCritic   # reuse the Version-3 model

def evaluate(ac, seed, episodes=10):
    env = gym.make("CartPole-v1")
    returns = []
    for i in range(episodes):
        obs, _ = env.reset(seed=seed * 100 + i)
        total = 0
        for _ in range(500):
            with torch.no_grad():
                act = ac.dist(torch.as_tensor(obs, dtype=torch.float32)).sample()
            obs, r, term, trunc, _ = env.step(act.item())
            total += r
            if term or trunc:
                break
        returns.append(total)
    env.close()
    return np.mean(returns), np.std(returns)

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--seed", type=int, default=0)
    parser.add_argument("--lr", type=float, default=3e-4)
    parser.add_argument("--total_steps", type=int, default=50_000)
    args = parser.parse_args()

    # Record the full config (the experiment's ID card)
    config = vars(args)
    torch.manual_seed(args.seed); np.random.seed(args.seed)

    ac = ActorCritic(4, 2)
    # ... plug in the Version-3 training loop here ...
    mean, std = evaluate(ac, args.seed)

    # Artifacts: model + config + metrics, all saved to disk
    torch.save(ac.state_dict(), f"model_seed{args.seed}.pt")
    with open(f"result_seed{args.seed}.json", "w", encoding="utf-8") as f:
        json.dump({**config, "eval_mean": mean, "eval_std": std}, f, ensure_ascii=False, indent=2)
    print(f"seed={args.seed} eval_return={mean:.1f} ± {std:.1f}")

if __name__ == "__main__":
    main()

Run three different seeds and merge the results:

bash
for s in 0 1 2; do python run_experiment.py --seed $s; done

That's the minimal closed loop — multi-seed, reproducible, artifacts on disk. It is exactly the smallest block of the experiment matrix in Evaluation in Practice. On top of it, add Hydra config management and W&B/TensorBoard logging, and you have an engineering-grade experiment system.

9. Extending to Other Environments ​

Once CartPole works, swap the environment name in gym.make to try harder tasks:

EnvironmentAction spaceDifficulty jumpWhat to change
LunarLander-v2Discrete, 4Sparser rewards, contact dynamicsAlmost nothing (all three versions handle discrete actions)
MountainCar-v0Discrete, 3Extremely sparse reward — plain Q-learning will struggleQ-learning needs reward shaping or a curriculum
Acrobot-v1Discrete, 3Continuous swing-up controlNo code changes; tune hyperparameters
Pendulum-v1Continuous, 1Your first continuous-action taskIn Version 3, swap Categorical for a Gaussian
HalfCheetah-v4/Humanoid-v4 (MuJoCo)Continuous, 6/17High-dimensional continuous controlNeeds vectorized sampling + GPU; a framework (SB3/Brax) is the pragmatic choice

Every new environment gets the same acceptance routine

A new environment counts as "working" when: the random baseline scores → a simple policy learns something → the three-way comparison table reproduces the "PPO is fast and steady" pattern. This is the minimal version of the "build baselines before swapping algorithms" practice emphasized in Evaluation in Practice.

For more environment libraries (Brax, MuJoCo, Procgen, Atari, and more), see the Datasets & Tools Index; for the algorithm families behind the three versions, see Value-Based Learning, Policy Gradient Methods, and The Actor-Critic Family.

Further Reading ​

References ​