Appearance
Tuning and Hyperparameter Optimization
In a nutshell: this page teaches a "diagnose first, then tune" methodology for RL hyperparameters — learning to read learning curves to locate the disease before touching the smallest set of knobs, in the right order; starting from defaults, through batch experiments (Optuna) and AutoRL. It solves the everyday predicament of "ran the defaults, no improvement, so I started twiddling everything at random." By the end you'll have a clear tuning decision tree.
First, one thing to internalize: RL is an order of magnitude more sensitive to hyperparameters than supervised learning. In supervised learning, learning_rate=1e-3 versus 1e-2 usually just changes how fast you converge; in RL, the same 1e-3 can mean the difference between "learns" and "collapses" depending on the implementation (PPO's GAE details, whether reward normalization is on). And RL adds a dimension supervised learning doesn't have: the same hyperparameters can behave wildly differently across seeds. Stack these two facts together and "tuning by feel" becomes a near-guaranteed losing strategy in RL.
This page and evaluation in practice are companion pieces: evaluation covers "how confident am I in an experiment's result," tuning covers "how to find a good configuration efficiently." Both share the same experiment infrastructure.
1. The Sensitivity Reality: Why RL Tuning Is So Hard
| Fact | Supervised learning | RL | Engineering implication |
|---|---|---|---|
| Sensitivity to learning rate | High (but wide workable range) | Extreme, and the workable range is narrow | Check the learning rate first |
| Stability of the training target | Static loss | Target drifts with the policy (non-stationary) | Curves can "improve, then collapse" |
| Evaluation noise | Stable validation set | High return variance + seed sensitivity | Must confirm across multiple seeds |
| Impact of implementation details | Some | Huge (reward normalization, done handling) | Reproduce first, tune second |
| Cost of one experiment | Minutes | Minutes to days | Build batch-experiment infrastructure early |
Two papers nail this down: Henderson et al. (2018) showed that RL algorithms' sensitivity to seeds and implementation details can mask algorithmic differences; Engstrom et al. (2020) ablated PPO directly and found that changing implementation details alone (learning-rate schedules, reward normalization) produces performance differences on the same order as "algorithmic improvements." So rule number one of tuning: make sure your experimental baseline is correct before tuning anything.
The most expensive trap: tuning on top of a buggy implementation
If your PPO forgot the GAE mask, skipped reward normalization, or treats truncated as failure — no amount of tuning will ever work, because the disease isn't in the hyperparameters. The first step in diagnosing a tuning problem is always "run a known benchmark (CartPole or the official default config)" to confirm the pipeline itself is sound. See the reproduction diagnostic checklist in common pitfalls.
2. Core Hyperparameters Cheat Sheet
Grouped by algorithm family, these are the main knobs that matter. "Sensitivity" here means "how much damage changing it does."
1. Universal (all algorithms)
| Parameter | Role | Typical range | Sensitivity | Tuning intuition |
|---|---|---|---|---|
learning_rate | Step size | 3e-4 ~ 3e-3 (PPO); 1e-4 ~ 1e-3 (SAC) | ★★★★★ | The prime suspect — check it first |
gamma (γ) | Discount factor | 0.9 ~ 0.999 | ★★★★ | The longer-horizon the task (delayed rewards), the closer γ to 1 |
network width/depth | Capacity | 64–512 wide, 2–3 layers | ★★★ | Too small underfits; too large is slow and prone to overfitting |
seed | Randomness | Fixed set | ★★★★ | Not a hyperparameter, but must be confirmed across seeds |
batch/minibatch size | Update granularity | 64–4096 | ★★★ | Large batches are stable but slow |
2. PPO-specific
| Parameter | Role | Typical range | Sensitivity | Tuning intuition |
|---|---|---|---|---|
n_steps | Sampling steps per update | 512–2048 (single env); by total steps after vectorization | ★★★★ | Too small → updates too frequent, unstable; too large → data goes stale |
gae_lambda (λ) | Bias/variance trade-off in advantage estimation | 0.9 ~ 0.99 | ★★★ | Near 1 leans Monte Carlo (high variance); near 0 leans TD |
clip_range (ε) | Maximum step size per update round | 0.1 ~ 0.3 | ★★★ | Larger means more aggressive updates |
update_epochs | Epochs reusing the same batch | 3 ~ 10 | ★★★ | Too many → overfitting to that batch |
ent_coef | Entropy regularization strength | 0 ~ 0.05 | ★★★★ | Turn up for insufficient exploration; turn down (or to zero) if training destabilizes or randomizes |
3. SAC/DDPG and other off-policy methods
| Parameter | Role | Typical range | Sensitivity | Tuning intuition |
|---|---|---|---|---|
buffer_size | Replay capacity | 1e5 ~ 1e6 | ★★ | Larger means more diverse but staler samples |
tau | Target network soft-update coefficient | 0.005 ~ 0.01 | ★★ | Smaller means slower target updates |
ent_coef / target_entropy | Entropy term | auto (auto) | ★★★ | Prefer automatic temperature — see the Actor-Critic family |
learning_starts | Replay warm-up steps | 1e3 ~ 1e4 | ★★ | Learning before the buffer fills causes oscillation |
Start from defaults, and don't touch these yet
SB3/CleanRL defaults are battle-tested (SB3's PPO defaults — lr=3e-4, n_steps=2048, clip=0.2 — are the RL community's agreed "stable starting point"). In the first round, change only one parameter and leave everything else at default. What you're tuning isn't "the best configuration" but "confirming the direction each knob moves things."
3. The Tuning Order: Reward → Environment → Algorithm → Hyperparameters
The biggest misconception in tuning is reaching for learning_rate first. The right priority is to start from the disease furthest upstream:
text
① Reward Sparse / miswritten / mis-scaled reward → no algorithm can help
(First check: average return magnitude, reward-term provenance — see reward engineering)
② Environment Env bug / observation missing information / wrong action space
(First check: environment smoke test, heuristic-baseline score)
③ Algorithm Wrong family chosen (needed off-policy, used on-policy)
(First check: sample budget, action space, stability requirements)
④ Hyperparameters Only when the first three are clean is it time to turn knobs
(start from the most sensitive in the section-2 tables)Each layer has a "one-sentence check":
| Layer | Check question | If something's off |
|---|---|---|
| Reward | How much return can a heuristic policy get? Is the reward bounded? | Fix the reward first, don't tune (see reward engineering) |
| Environment | Does the random-policy smoke test pass? Are observation dimensions correct? | Fix the environment first |
| Algorithm | Is the sample budget enough for off-policy to shine? Does the action space match? | Switch algorithm family first (see the Actor-Critic family) |
| Hyperparameters | Which shape is the learning curve? | Apply the right cure from the section-4 diagnostic table |
Why the order matters so much
A real case: a team moved PPO's lr from 3e-4 to 3e-3 and clip from 0.2 to 0.1, spent two weeks going nowhere — and finally discovered the reward function was missing a negative term, so every policy was stuck wandering a plateau of identical scores. Tuning hyperparameters before the reward is fixed is like re-tiling a building whose foundation is tilted.
4. The Diagnostic Tool: Learning-Curve Shape Analysis
"Tuning the right cure for the right disease" starts with reading the curve. The five shapes below are the most common "diseases" in RL tuning, each with a clear prescription:
text
① Flatlines ② Rises, then collapses ③ Violent oscillation
return return return
│ │┐ │╱╲╱╲╱╲
│ ───────── │╱└──┐ │╱╲╱╲╱╲╱
│ ▁▁▁▁▁ │ ╲╲ │╱╲╱╲╱╲
└──────────────steps └────────steps └────────steps
④ Slow climb (no convergence) ⑤ Early surge, then plateau
return return
│ ───── │ ────────
│ ╱ │ ╱
│ ╱ │ ╱
│╱ │╱
└────────────steps └────────────steps| Shape | Primary cause | First prescription | Backup |
|---|---|---|---|
| ① Flatlines | Reward/environment problem, or learning rate too small | Go back to section 3 and check reward + environment | Bump lr up one notch |
| ② Rises, then collapses | Learning rate too high / clip too large / entropy too small | Drop lr one notch (for PPO also drop clip) | Add gradient clipping |
| ③ Violent oscillation | High variance (on-policy + small batch) | Increase n_steps/batch | Increase gae_lambda or entropy regularization |
| ④ Slow climb | Learning rate too small / entropy too large / network too small | Raise lr | Widen the network |
| ⑤ Early surge, then plateau | Stuck at a local optimum / insufficient exploration | Increase entropy regularization or initial exploration | Switch seeds to confirm variance vs. a true plateau |
Curve diagnosis must be done across multiple seeds
A single curve that "rises then collapses" might be a genuine collapse — or just a bad-luck seed. Shape judgments need at least 3 overlapping seed curves; only diagnose when the shapes agree. This is also why evaluation in practice makes the IQR band the default view — the IQR band is the best diagnostic chart you can get.
Supplementary diagnostics: more than just return
Return is the final outcome, but intermediate metrics can tell you where the disease sits before it shows up in return:
| Diagnostic | Healthy shape | Abnormal shape → likely cause |
|---|---|---|
| Policy entropy | Gradual decline | Plunge → premature convergence; never drops → not learning |
| Value loss / TD error | Gradual decline | Wild swings → learning rate too high or unstable bootstrapping |
| Probability ratio (PPO) | Clustered near 1 | Far from 1 → clip is firing constantly, updates too aggressive |
| Average return | Rising | See the table above |
5. Random Search vs. Bayesian Optimization (Optuna)
When point-wise "right cure for the right disease" tuning hits a wall, or you need to explore an unfamiliar configuration space, move to batch search.
| Method | Principle | Pros | Cons | When to use |
|---|---|---|---|---|
| Grid search | Full cross-product of a few values per parameter | Simple, interpretable | Combinatorial explosion, wasteful | Very few parameters |
| Random search | Randomly sample N configs from the ranges | More efficient than grid (especially in high dimensions), simple | Doesn't exploit past results | The default starting point |
| Bayesian optimization (Optuna/TPE) | Build a probabilistic model from past results and pick the most promising next point | Sample-efficient, automatic pruning | Higher implementation complexity, serial single points | Limited budget, second-round refinement |
Practical playbook:
- Reconnoiter with random search first: run 30–50 random configs to establish "what range of configs can learn at all," and get a rough upper bound along the way.
- Finish with Bayesian optimization: feed the random-search results to Optuna and let it mine the good region.
- Record every search: store each config + result in a table. When the search ends, you walk away with more than the best config — you get "which parameter matters most" (Optuna's
importanceanalysis).
A minimal Optuna example (PPO-CartPole)
python
import optuna
from stable_baselines3 import PPO
from stable_baselines3.common.env_util import make_vec_env
def objective(trial):
# Sample one set of hyperparameters per trial
lr = trial.suggest_float("learning_rate", 1e-4, 1e-2, log=True)
gae = trial.suggest_float("gae_lambda", 0.9, 0.99)
ent = trial.suggest_float("ent_coef", 0.0, 0.05)
env = make_vec_env("CartPole-v1", n_envs=4)
model = PPO("MlpPolicy", env, learning_rate=lr,
gae_lambda=gae, ent_coef=ent, seed=0)
model.learn(total_timesteps=20_000)
# Evaluate (use fixed evaluation seeds — not the training trajectories)
eval_env = make_vec_env("CartPole-v1", n_envs=1)
obs = eval_env.reset()
total = 0.0
for _ in range(500):
action, _ = model.predict(obs, deterministic=True)
obs, r, done, info = eval_env.step(action)
total += r.item()
if done.all():
break
return total # Optuna maximizes this
study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=30)
print("best params:", study.best_params)Optuna's hidden traps
- The fixed-seed trap: the example above pins
seed=0— if that seed happens to favor some parameter region, Optuna will chase luck. The proper approach is 3 seeds per config, using the median as the objective. - Budget consistency:
total_timestepsmust be identical across trials, otherwise you're comparing budgets, not parameters. - Prune with care: Optuna's early stopping cuts configs that look bad early, but some RL configs are precisely the "bad at first, great later" kind (high entropy, long exploration). Prune cautiously.
6. The Fixed-Seed Trap and Multi-Seed Confirmation
This is the most brutally punishing rule in RL tuning. The full background is in common pitfalls and anti-patterns; here are just the operating rules:
| Scenario | Allowed seed usage | Not allowed |
|---|---|---|
| Quick parameter scouting | Single seed to eyeball the shape | Treating a single-seed result as a final conclusion |
| Comparing two configurations | The same seed set, paired comparisons | Config A on seeds 0–4, config B on seeds 5–9 |
| Final conclusions | 3–5+ seeds, median ± IQR | Reporting "the best run we had" |
The sneakiest trap: tuning against the validation config
If you run a batch of configs, pick the one that did best at seed=0, and then "validate" it at seed=0 — you've already overfit to that same randomness. Validation must use a fresh seed set. That's seed overfitting: your "tuning" is really memorizing a set of random numbers.
7. A Brief Intro to AutoRL
AutoRL tries to automate the whole "algorithm + hyperparameters + architecture" search — essentially wrapping the RL training loop in an outer layer of meta-learning or evolution. A few directions:
| Direction | Approach | Representative work | Maturity |
|---|---|---|---|
| Hyperparameter search | Batch experiments with Optuna/BOHB etc. | SB3 + Optuna tutorials | Mature, production-ready |
| Dynamic tuning during training | Adjust lr/entropy on the fly based on metrics | PBT (Population Based Training) | Common in papers, rare in practice |
| Automatic objective design | Let an algorithm learn update rules / objective functions | Learned Policy Gradient, Meta-RL | Research frontier |
Judgments from an engineering standpoint:
- PBT (population-based training) is the best deal in AutoRL: a population of runs trains in parallel; underperformers copy the weights and hyperparameters of strong performers, then mutate. OpenAI used it to set SOTA on several benchmarks (Lilian Weng's blog has a detailed write-up).
- Don't expect AutoRL to replace understanding: search just outsources "tuning craft" to compute — and it needs you to define the search space, objective, and evaluation protocol correctly first, which is exactly what the earlier sections teach.
The boundary of AutoRL and search
The AutoRL you imagine: "give it an environment, it automatically finds an algorithm that hits SOTA." The AutoRL you get: "you define the search space, evaluation protocol, budget, and seed scheme, and it helps you exhaust them." The former doesn't exist; the latter is hugely valuable. The pragmatic path: manual diagnosis + random search first, Optuna to finish; consider PBT only when the team has distributed infrastructure.
8. Experiment Management: Configs, Logs, Reproducibility
The ultimate determinant of tuning efficiency isn't "how fast you tune" but "not wasting a single experiment." A traceable experiment system is what makes large-scale search affordable — otherwise every tuning session is a blind box.
1. Config as Code: Hydra
yaml
# config.yaml (the complete ID card of one experiment)
seed: 42
env_id: HalfCheetah-v4
algorithm:
name: ppo
learning_rate: 3.0e-4
gamma: 0.99
gae_lambda: 0.95
clip_range: 0.2
ent_coef: 0.0
search: # if this is a batch search, record the search method
method: random
n_trials: 40With Hydra or a simple YAML loader, override any field from the command line (python train.py algorithm.learning_rate=1e-3), and archive the full config automatically with every experiment — "reproduction" degrades to "re-run a directory."
2. Logging and Versioning
- Log during training: every N steps, write one row of
(step, train_return, eval_return, entropy, loss, lr)to CSV or W&B. - Code version: record the
git commitin the experiment directory. - Dependency versions: record
pip freezeor the lockfile.
3. Minimum Discipline for One Tuning Session
text
Every experiment round (recommended as a checklist):
□ Change only one variable (or explicitly record which ones changed)
□ Record the before/after hyperparameters and results
□ Confirm the shape on at least 3 seeds
□ Write the conclusion into the experiment log (even one line: "lr 3e-4→1e-3 collapsed, shape ③")"Change one variable at a time" is discipline, not dogma
Batch search (Optuna) changing many variables at once is fine — that's systematic exploration. What you must avoid is random-style tuning where you hand-tweak five parameters at once and can't say why. Discipline = every change is recorded, every conclusion is backed.
9. Tuning Decision Tree: Summary
text
Policy not learning?
├─ Ran a known benchmark (CartPole / official defaults) to confirm the pipeline is OK? ── No → Fix the implementation, don't tune
├─ Reward problem? ── Yes → Fix the reward
├─ Environment problem? ── Yes → Fix the environment
├─ Wrong algorithm family? ── Yes → Switch algorithms
└─ Read the learning-curve shape
├─ Flatlines → raise lr
├─ Rises, collapses → lower lr / clip, add gradient clipping
├─ Oscillates → increase n_steps / batch
├─ Climbs slowly → raise lr / lower entropy
└─ Plateaus → raise entropy / switch seeds to confirm
Then: multi-seed confirmation → batch search (random → Optuna) → record + reportFurther Reading
- Policy gradient methods — where PPO hyperparameters (clip, entropy, GAE) come from
- The Actor-Critic family — the theory behind off-policy hyperparameters: SAC automatic temperature, TD3 double-Q, etc.
- Build an RL evaluation setup from scratch — the shared experiment infrastructure and statistical rigor behind tuning and evaluation
- Common pitfalls and anti-patterns — seed overfitting, compute illusions, and other wrecks on the tuning road
- How to choose frameworks and tools — selecting the SB3/CleanRL/Optuna toolchain
References
- Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
- Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292
- Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133
- Bergstra, J. & Bengio, Y. (2012). Random Search for Hyper-Parameter Optimization. JMLR 13, 281–305. (the classic argument that random search beats grid search)
- Optuna documentation: optuna.org; the SB3 Optuna integration example is at github.com/DLR-RM/rl-baselines3-zoo
- Jaderberg, M. et al. (2017). Population Based Training of Neural Networks. arXiv:1711.09846 (the PBT paper)
- Lilian Weng (2019). A Gentle Introduction to PBT / AutoRL survey. lilianweng.github.io
- Hydra documentation: hydra.cc