Skip to content

Tuning and Hyperparameter Optimization

On this page RL is brutally sensitive to hyperparameters — learning rate, γ, GAE λ, entropy coefficient, clip range, network width; the right order to tune from defaults; AutoRL and batch experiments; and a "diagnose before you tune" methodology.

Tuning and Hyperparameter Optimization ​

In a nutshell: this page teaches a "diagnose first, then tune" methodology for RL hyperparameters — learning to read learning curves to locate the disease before touching the smallest set of knobs, in the right order; starting from defaults, through batch experiments (Optuna) and AutoRL. It solves the everyday predicament of "ran the defaults, no improvement, so I started twiddling everything at random." By the end you'll have a clear tuning decision tree.

First, one thing to internalize: RL is an order of magnitude more sensitive to hyperparameters than supervised learning. In supervised learning, learning_rate=1e-3 versus 1e-2 usually just changes how fast you converge; in RL, the same 1e-3 can mean the difference between "learns" and "collapses" depending on the implementation (PPO's GAE details, whether reward normalization is on). And RL adds a dimension supervised learning doesn't have: the same hyperparameters can behave wildly differently across seeds. Stack these two facts together and "tuning by feel" becomes a near-guaranteed losing strategy in RL.

This page and evaluation in practice are companion pieces: evaluation covers "how confident am I in an experiment's result," tuning covers "how to find a good configuration efficiently." Both share the same experiment infrastructure.

1. The Sensitivity Reality: Why RL Tuning Is So Hard ​

FactSupervised learningRLEngineering implication
Sensitivity to learning rateHigh (but wide workable range)Extreme, and the workable range is narrowCheck the learning rate first
Stability of the training targetStatic lossTarget drifts with the policy (non-stationary)Curves can "improve, then collapse"
Evaluation noiseStable validation setHigh return variance + seed sensitivityMust confirm across multiple seeds
Impact of implementation detailsSomeHuge (reward normalization, done handling)Reproduce first, tune second
Cost of one experimentMinutesMinutes to daysBuild batch-experiment infrastructure early

Two papers nail this down: Henderson et al. (2018) showed that RL algorithms' sensitivity to seeds and implementation details can mask algorithmic differences; Engstrom et al. (2020) ablated PPO directly and found that changing implementation details alone (learning-rate schedules, reward normalization) produces performance differences on the same order as "algorithmic improvements." So rule number one of tuning: make sure your experimental baseline is correct before tuning anything.

The most expensive trap: tuning on top of a buggy implementation

If your PPO forgot the GAE mask, skipped reward normalization, or treats truncated as failure — no amount of tuning will ever work, because the disease isn't in the hyperparameters. The first step in diagnosing a tuning problem is always "run a known benchmark (CartPole or the official default config)" to confirm the pipeline itself is sound. See the reproduction diagnostic checklist in common pitfalls.

2. Core Hyperparameters Cheat Sheet ​

Grouped by algorithm family, these are the main knobs that matter. "Sensitivity" here means "how much damage changing it does."

1. Universal (all algorithms) ​

ParameterRoleTypical rangeSensitivityTuning intuition
learning_rateStep size3e-4 ~ 3e-3 (PPO); 1e-4 ~ 1e-3 (SAC)★★★★★The prime suspect — check it first
gamma (γ)Discount factor0.9 ~ 0.999★★★★The longer-horizon the task (delayed rewards), the closer γ to 1
network width/depthCapacity64–512 wide, 2–3 layers★★★Too small underfits; too large is slow and prone to overfitting
seedRandomnessFixed set★★★★Not a hyperparameter, but must be confirmed across seeds
batch/minibatch sizeUpdate granularity64–4096★★★Large batches are stable but slow

2. PPO-specific ​

ParameterRoleTypical rangeSensitivityTuning intuition
n_stepsSampling steps per update512–2048 (single env); by total steps after vectorization★★★★Too small → updates too frequent, unstable; too large → data goes stale
gae_lambda (λ)Bias/variance trade-off in advantage estimation0.9 ~ 0.99★★★Near 1 leans Monte Carlo (high variance); near 0 leans TD
clip_range (ε)Maximum step size per update round0.1 ~ 0.3★★★Larger means more aggressive updates
update_epochsEpochs reusing the same batch3 ~ 10★★★Too many → overfitting to that batch
ent_coefEntropy regularization strength0 ~ 0.05★★★★Turn up for insufficient exploration; turn down (or to zero) if training destabilizes or randomizes

3. SAC/DDPG and other off-policy methods ​

ParameterRoleTypical rangeSensitivityTuning intuition
buffer_sizeReplay capacity1e5 ~ 1e6★★Larger means more diverse but staler samples
tauTarget network soft-update coefficient0.005 ~ 0.01★★Smaller means slower target updates
ent_coef / target_entropyEntropy termauto (auto)★★★Prefer automatic temperature — see the Actor-Critic family
learning_startsReplay warm-up steps1e3 ~ 1e4★★Learning before the buffer fills causes oscillation

Start from defaults, and don't touch these yet

SB3/CleanRL defaults are battle-tested (SB3's PPO defaults — lr=3e-4, n_steps=2048, clip=0.2 — are the RL community's agreed "stable starting point"). In the first round, change only one parameter and leave everything else at default. What you're tuning isn't "the best configuration" but "confirming the direction each knob moves things."

3. The Tuning Order: Reward → Environment → Algorithm → Hyperparameters ​

The biggest misconception in tuning is reaching for learning_rate first. The right priority is to start from the disease furthest upstream:

text
① Reward        Sparse / miswritten / mis-scaled reward → no algorithm can help
                   (First check: average return magnitude, reward-term provenance — see reward engineering)
② Environment   Env bug / observation missing information / wrong action space
                   (First check: environment smoke test, heuristic-baseline score)
③ Algorithm     Wrong family chosen (needed off-policy, used on-policy)
                   (First check: sample budget, action space, stability requirements)
④ Hyperparameters   Only when the first three are clean is it time to turn knobs
                       (start from the most sensitive in the section-2 tables)

Each layer has a "one-sentence check":

LayerCheck questionIf something's off
RewardHow much return can a heuristic policy get? Is the reward bounded?Fix the reward first, don't tune (see reward engineering)
EnvironmentDoes the random-policy smoke test pass? Are observation dimensions correct?Fix the environment first
AlgorithmIs the sample budget enough for off-policy to shine? Does the action space match?Switch algorithm family first (see the Actor-Critic family)
HyperparametersWhich shape is the learning curve?Apply the right cure from the section-4 diagnostic table

Why the order matters so much

A real case: a team moved PPO's lr from 3e-4 to 3e-3 and clip from 0.2 to 0.1, spent two weeks going nowhere — and finally discovered the reward function was missing a negative term, so every policy was stuck wandering a plateau of identical scores. Tuning hyperparameters before the reward is fixed is like re-tiling a building whose foundation is tilted.

4. The Diagnostic Tool: Learning-Curve Shape Analysis ​

"Tuning the right cure for the right disease" starts with reading the curve. The five shapes below are the most common "diseases" in RL tuning, each with a clear prescription:

text
① Flatlines          ② Rises, then collapses    ③ Violent oscillation
return               return                     return
 │                     │┐                         │╱╲╱╲╱╲
 │   ─────────          │╱└──┐                    │╱╲╱╲╱╲╱
 │   ▁▁▁▁▁             │   ╲╲                    │╱╲╱╲╱╲
 └──────────────steps  └────────steps             └────────steps

④ Slow climb (no convergence)   ⑤ Early surge, then plateau
return                          return
 │      ─────                   │    ────────
 │    ╱                         │   ╱
 │  ╱                           │ ╱
 │╱                             │╱
 └────────────steps             └────────────steps
ShapePrimary causeFirst prescriptionBackup
① FlatlinesReward/environment problem, or learning rate too smallGo back to section 3 and check reward + environmentBump lr up one notch
② Rises, then collapsesLearning rate too high / clip too large / entropy too smallDrop lr one notch (for PPO also drop clip)Add gradient clipping
③ Violent oscillationHigh variance (on-policy + small batch)Increase n_steps/batchIncrease gae_lambda or entropy regularization
④ Slow climbLearning rate too small / entropy too large / network too smallRaise lrWiden the network
⑤ Early surge, then plateauStuck at a local optimum / insufficient explorationIncrease entropy regularization or initial explorationSwitch seeds to confirm variance vs. a true plateau

Curve diagnosis must be done across multiple seeds

A single curve that "rises then collapses" might be a genuine collapse — or just a bad-luck seed. Shape judgments need at least 3 overlapping seed curves; only diagnose when the shapes agree. This is also why evaluation in practice makes the IQR band the default view — the IQR band is the best diagnostic chart you can get.

Supplementary diagnostics: more than just return ​

Return is the final outcome, but intermediate metrics can tell you where the disease sits before it shows up in return:

DiagnosticHealthy shapeAbnormal shape → likely cause
Policy entropyGradual declinePlunge → premature convergence; never drops → not learning
Value loss / TD errorGradual declineWild swings → learning rate too high or unstable bootstrapping
Probability ratio (PPO)Clustered near 1Far from 1 → clip is firing constantly, updates too aggressive
Average returnRisingSee the table above

5. Random Search vs. Bayesian Optimization (Optuna) ​

When point-wise "right cure for the right disease" tuning hits a wall, or you need to explore an unfamiliar configuration space, move to batch search.

MethodPrincipleProsConsWhen to use
Grid searchFull cross-product of a few values per parameterSimple, interpretableCombinatorial explosion, wastefulVery few parameters
Random searchRandomly sample N configs from the rangesMore efficient than grid (especially in high dimensions), simpleDoesn't exploit past resultsThe default starting point
Bayesian optimization (Optuna/TPE)Build a probabilistic model from past results and pick the most promising next pointSample-efficient, automatic pruningHigher implementation complexity, serial single pointsLimited budget, second-round refinement

Practical playbook:

  1. Reconnoiter with random search first: run 30–50 random configs to establish "what range of configs can learn at all," and get a rough upper bound along the way.
  2. Finish with Bayesian optimization: feed the random-search results to Optuna and let it mine the good region.
  3. Record every search: store each config + result in a table. When the search ends, you walk away with more than the best config — you get "which parameter matters most" (Optuna's importance analysis).

A minimal Optuna example (PPO-CartPole) ​

python
import optuna
from stable_baselines3 import PPO
from stable_baselines3.common.env_util import make_vec_env

def objective(trial):
    # Sample one set of hyperparameters per trial
    lr = trial.suggest_float("learning_rate", 1e-4, 1e-2, log=True)
    gae = trial.suggest_float("gae_lambda", 0.9, 0.99)
    ent = trial.suggest_float("ent_coef", 0.0, 0.05)
    env = make_vec_env("CartPole-v1", n_envs=4)
    model = PPO("MlpPolicy", env, learning_rate=lr,
                gae_lambda=gae, ent_coef=ent, seed=0)
    model.learn(total_timesteps=20_000)
    # Evaluate (use fixed evaluation seeds — not the training trajectories)
    eval_env = make_vec_env("CartPole-v1", n_envs=1)
    obs = eval_env.reset()
    total = 0.0
    for _ in range(500):
        action, _ = model.predict(obs, deterministic=True)
        obs, r, done, info = eval_env.step(action)
        total += r.item()
        if done.all():
            break
    return total   # Optuna maximizes this

study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=30)
print("best params:", study.best_params)

Optuna's hidden traps

  1. The fixed-seed trap: the example above pins seed=0 — if that seed happens to favor some parameter region, Optuna will chase luck. The proper approach is 3 seeds per config, using the median as the objective.
  2. Budget consistency: total_timesteps must be identical across trials, otherwise you're comparing budgets, not parameters.
  3. Prune with care: Optuna's early stopping cuts configs that look bad early, but some RL configs are precisely the "bad at first, great later" kind (high entropy, long exploration). Prune cautiously.

6. The Fixed-Seed Trap and Multi-Seed Confirmation ​

This is the most brutally punishing rule in RL tuning. The full background is in common pitfalls and anti-patterns; here are just the operating rules:

ScenarioAllowed seed usageNot allowed
Quick parameter scoutingSingle seed to eyeball the shapeTreating a single-seed result as a final conclusion
Comparing two configurationsThe same seed set, paired comparisonsConfig A on seeds 0–4, config B on seeds 5–9
Final conclusions3–5+ seeds, median ± IQRReporting "the best run we had"

The sneakiest trap: tuning against the validation config

If you run a batch of configs, pick the one that did best at seed=0, and then "validate" it at seed=0 — you've already overfit to that same randomness. Validation must use a fresh seed set. That's seed overfitting: your "tuning" is really memorizing a set of random numbers.

7. A Brief Intro to AutoRL ​

AutoRL tries to automate the whole "algorithm + hyperparameters + architecture" search — essentially wrapping the RL training loop in an outer layer of meta-learning or evolution. A few directions:

DirectionApproachRepresentative workMaturity
Hyperparameter searchBatch experiments with Optuna/BOHB etc.SB3 + Optuna tutorialsMature, production-ready
Dynamic tuning during trainingAdjust lr/entropy on the fly based on metricsPBT (Population Based Training)Common in papers, rare in practice
Automatic objective designLet an algorithm learn update rules / objective functionsLearned Policy Gradient, Meta-RLResearch frontier

Judgments from an engineering standpoint:

  • PBT (population-based training) is the best deal in AutoRL: a population of runs trains in parallel; underperformers copy the weights and hyperparameters of strong performers, then mutate. OpenAI used it to set SOTA on several benchmarks (Lilian Weng's blog has a detailed write-up).
  • Don't expect AutoRL to replace understanding: search just outsources "tuning craft" to compute — and it needs you to define the search space, objective, and evaluation protocol correctly first, which is exactly what the earlier sections teach.

The boundary of AutoRL and search

The AutoRL you imagine: "give it an environment, it automatically finds an algorithm that hits SOTA." The AutoRL you get: "you define the search space, evaluation protocol, budget, and seed scheme, and it helps you exhaust them." The former doesn't exist; the latter is hugely valuable. The pragmatic path: manual diagnosis + random search first, Optuna to finish; consider PBT only when the team has distributed infrastructure.

8. Experiment Management: Configs, Logs, Reproducibility ​

The ultimate determinant of tuning efficiency isn't "how fast you tune" but "not wasting a single experiment." A traceable experiment system is what makes large-scale search affordable — otherwise every tuning session is a blind box.

1. Config as Code: Hydra ​

yaml
# config.yaml (the complete ID card of one experiment)
seed: 42
env_id: HalfCheetah-v4

algorithm:
  name: ppo
  learning_rate: 3.0e-4
  gamma: 0.99
  gae_lambda: 0.95
  clip_range: 0.2
  ent_coef: 0.0

search:            # if this is a batch search, record the search method
  method: random
  n_trials: 40

With Hydra or a simple YAML loader, override any field from the command line (python train.py algorithm.learning_rate=1e-3), and archive the full config automatically with every experiment — "reproduction" degrades to "re-run a directory."

2. Logging and Versioning ​

  • Log during training: every N steps, write one row of (step, train_return, eval_return, entropy, loss, lr) to CSV or W&B.
  • Code version: record the git commit in the experiment directory.
  • Dependency versions: record pip freeze or the lockfile.

3. Minimum Discipline for One Tuning Session ​

text
Every experiment round (recommended as a checklist):
  □ Change only one variable (or explicitly record which ones changed)
  □ Record the before/after hyperparameters and results
  □ Confirm the shape on at least 3 seeds
  □ Write the conclusion into the experiment log (even one line: "lr 3e-4→1e-3 collapsed, shape ③")

"Change one variable at a time" is discipline, not dogma

Batch search (Optuna) changing many variables at once is fine — that's systematic exploration. What you must avoid is random-style tuning where you hand-tweak five parameters at once and can't say why. Discipline = every change is recorded, every conclusion is backed.

9. Tuning Decision Tree: Summary ​

text
Policy not learning?
  ├─ Ran a known benchmark (CartPole / official defaults) to confirm the pipeline is OK? ── No → Fix the implementation, don't tune
  ├─ Reward problem?                                          ── Yes → Fix the reward
  ├─ Environment problem?                                     ── Yes → Fix the environment
  ├─ Wrong algorithm family?                                  ── Yes → Switch algorithms
  └─ Read the learning-curve shape
       ├─ Flatlines    → raise lr
       ├─ Rises, collapses → lower lr / clip, add gradient clipping
       ├─ Oscillates   → increase n_steps / batch
       ├─ Climbs slowly → raise lr / lower entropy
       └─ Plateaus     → raise entropy / switch seeds to confirm
    Then: multi-seed confirmation → batch search (random → Optuna) → record + report

Further Reading ​

References ​

  • Henderson, P. et al. (2018). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560
  • Engstrom, L. et al. (2020). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. ICLR 2020. arXiv:1805.08292
  • Islam, R. et al. (2017). Reproducibility of Benchmarked Deep Reinforcement Learning Tasks. arXiv:1708.04133
  • Bergstra, J. & Bengio, Y. (2012). Random Search for Hyper-Parameter Optimization. JMLR 13, 281–305. (the classic argument that random search beats grid search)
  • Optuna documentation: optuna.org; the SB3 Optuna integration example is at github.com/DLR-RM/rl-baselines3-zoo
  • Jaderberg, M. et al. (2017). Population Based Training of Neural Networks. arXiv:1711.09846 (the PBT paper)
  • Lilian Weng (2019). A Gentle Introduction to PBT / AutoRL survey. lilianweng.github.io
  • Hydra documentation: hydra.cc