Theme
Tuning Practice and Hyperparameter Optimization
Every minute you spend on tuning should first answer one question: is this parameter really worth tuning?
Hyperparameter optimization (HPO) is the most paradoxical part of machine learning engineering: it's described as critically important — "tuning is a craft," "the gap between experts and beginners comes down to this" — but the reality is that 90% of tuning in projects wastes compute, and what truly determines a model's ceiling is usually the data, features, and model choice. The goal of this article is not to teach you to memorize parameter values, but to help you build a complete methodology for "what to tune, when to tune it, what method to use, and how to ensure you're tuning correctly."
Let's start with a counterintuitive fact: a widely circulated Kaggle post found that top contestants spend far more time on feature engineering and data cleaning than on hyperparameter tuning — and when they do tune, they rarely use brute-force grid search, preferring instead "a few parameters, quick validation, narrowing by experience." This isn't because they've memorized parameter values, but because they put tuning in the right place.
The article is structured as follows: first, the tuning mindset (when you should actually tune); then a panoramic map of hyperparameters (what key knobs exist for each model family); then a comparison of four search methods; followed by complete code walkthroughs with sklearn and Optuna; then a deep dive into deep learning tuning; and finally, a set of tuning discipline rules to ensure your experiments are trustworthy.
I. Tuning Mindset: Data First, Then Model, Then Hyperparameters
1. Why "tuning from the start" is the biggest waste
The most common path for a beginner given a modeling task is: install XGBoost → run GridSearchCV overnight → stare at best_params_ with excitement → submit. The problem with this workflow isn't the tool, it's the order: hyperparameters are the last knob on the model, and four layers of problems lie above it that need to be investigated first.
┌─────────────────────────────────────────────────────────────┐
│ Problem-investigation order (top to bottom, higher = prior) │
├─────────────────────────────────────────────────────────────┤
│ ① Data layer Enough samples? Labels correct? Leakage? │
│ Normal distribution? │
│ ② Feature layer Any useful features left unused? Missing/ │
│ outlier handling? │
│ ③ Model layer Right model family? Bias or variance issue? │
│ ④ Hyperparam Only after ①–③ are settled do hyperparams │
│ get their turn │
└─────────────────────────────────────────────────────────────┘Why this order? Because higher-layer problems corrupt the signals at lower layers. If your data has leakage (e.g., deduplication before splitting), every "optimal value" you see at the hyperparameter level is fake — the model is just memorizing the answers, and you're painstakingly tuning the mouth shape for "memorizing the answers." For a complete list of fatal traps like data leakage, temporal leakage, and train/online inconsistency, see Common Pitfalls and Anti-Patterns; for a complete project workflow framework, see Building an ML Project from Scratch.
Don't tune on a shaky foundation
If you see high training metrics but poor validation performance — don't doubt your hyperparameters first, doubt overfitting and data leakage; if even the training metrics are low — don't rush to tune, first check whether your features and model match the problem. Hyperparameters can only amplify your gains once the foundation is solid; when the foundation is skewed, they only amplify errors.
2. Define "what tuning can solve" first
Hyperparameters can only affect one thing: the specific landing point of the bias-variance trade-off. Whether the model is underfitting (high bias) or overfitting (high variance) determines which direction you should turn:
| Symptom | Diagnosis | Tuning Direction |
|---|---|---|
| Training set & validation set both fail | Underfitting / high bias | Increase capacity: deeper models/more trees, reduce regularization, lower min_samples_leaf |
| Training set perfect, validation set poor | Overfitting / high variance | Strengthen regularization: add dropout/weight_decay, increase min_child_weight, reduce max_depth |
| Training set & validation set both near theoretical lower bound | Reached model ceiling | Stop tuning, go back to data and feature levels for gains |
This judgment method is also called "bias-variance decomposition," the most practical engineering technique in Model Evaluation and Validation. Before any tuning, plot the training and validation errors against model capacity — you'll immediately know which segment of the curve you're on.
3. When hyperparameters actually get their turn
Check the following checklist before entering the tuning stage — only proceed when all conditions are met:
- A baseline model is running: a simple logistic regression or single decision tree already has reasonable performance, and you know what the "blind lower bound" is;
- Data and features are frozen: no more frequent adding/removing of features — otherwise, every previous tuning result is invalidated by a feature change;
- Evaluation protocol is fixed: splitting method, evaluation metrics, and cross-validation folds are all set in stone, see Building a Model Evaluation Pipeline from Scratch;
- Goal is clear: precision-prioritized (competitions, offline tasks) or latency/resource-prioritized (online inference). This determines the direction of the search space.
Tuning before meeting these conditions is like framing a painting on a wobbly easel.
II. Hyperparameter Panorama
1. Parameters vs. Hyperparameters
First, distinguish two concepts: parameters (parameters) are learned by the model from the data, such as coefficients w, b in linear regression or weights in neural networks — their values are determined by the training algorithm and don't need your intervention. Hyperparameters (hyperparameters) are the knobs you set before training, such as the number of trees, learning rate, and regularization strength — the data won't tell you automatically what values they should take, and that's exactly what tuning solves.
Parameters: exist only after fit(X, y) → determined by the optimization algorithm
Hyperparams: exist before fit(X, y) → determined by the "tuner"A practical rule: the fewer hyperparameters a model has, and the less sensitive it is to them, the easier it is to use; the stronger the interactions between hyperparameters, the more systematic search is needed. Random Forest has almost only one "any value works" parameter (n_estimators), which is why it's the universal baseline for Kaggle. Deep learning has seven or eight interdependent knobs, which is why it requires specialized tuning techniques.
2. Tree Models and Ensemble Methods
The tree model family (decision trees, random forests, XGBoost, LightGBM, CatBoost) is currently the de facto standard for tabular data; see Tree Models and Ensemble Learning for detailed mechanics. Their hyperparameters can be divided into four groups: structure (how deep each tree grows), ensemble (how many trees, how they sample), regularization (prevent overfitting), and training (learning rate, etc.).
| Hyperparameter | Meaning | Direction | Common Range | Tuning Tips |
|---|---|---|---|---|
n_estimators / num_boost_round | Number of trees | More = lower bias, but diminishing returns, linear time cost | 100–2000 | Tune last; early stopping is better than stacking trees |
max_depth | Maximum depth of a single tree | Deeper = stronger fit, more overfitting | 3–15 | The #1 knob for controlling overfitting |
learning_rate / eta | Step size contributed by each tree | Smaller = more stable, needs more trees | 0.01–0.3 | Coupled with tree count: halve lr, often need to double tree count |
min_samples_split / min_child_weight | Minimum samples/weight for splitting | Larger = more overfit prevention | 2–20 / 1–10 | Prioritize increasing when overfitting |
min_samples_leaf | Minimum samples per leaf | Larger = smoother trees | 1–20 | Closely related to small-sample datasets |
max_features | Number of features considered per split | Smaller = lower variance | sqrt or 0.3–0.7 | Effective to reduce when features are many |
subsample / bagging_fraction | Sample ratio per tree | Smaller = more overfit prevention | 0.6–1.0 | Common in GBDT; default for random forest |
colsample_bytree | Feature ratio used per tree | Same as above | 0.6–1.0 | XGBoost-specific, similar to max_features |
lambda / reg_lambda | L2 regularization on leaf weights | Larger = more overfit prevention | 1–100 | Recommended when many trees, large depth |
Tree model tuning order (time-saving golden sequence)
- Fix learning rate (default 0.1) and enough trees, tune
max_depthandmin_child_weightto a "underfit/overfit boundary"; - Then tune sampling params (subsample, colsample_bytree, max_features);
- Finally tune regularization (lambda) with early stopping to find optimal tree count. This order decouples interdependent parameters, moving only one or two knobs at each step.
3. Linear Models and Regularization
Linear models (linear regression, Ridge, Lasso, ElasticNet, logistic regression, linear SVM) have few but precise hyperparameters, almost all centered around "regularization strength". See Overfitting and Regularization for their mechanics and regularization mechanisms.
| Hyperparameter | Meaning | Common Range | Notes |
|---|---|---|---|
C (logistic regression, LinearSVC) | Inverse of regularization strength: larger C = weaker penalty | 1e-4–1e4 (log grid) | Easiest pitfall: C direction is opposite to "strength" |
alpha (Ridge/Lasso/ElasticNet) | Regularization strength: larger = stronger penalty | 1e-4–10 | Opposite direction to C, be careful |
l1_ratio (ElasticNet) | Mix ratio of L1 to L2 | 0–1 | 0 = pure Ridge, 1 = pure Lasso |
penalty / solver | Regularization type and solver algorithm | See sklearn docs | Usually set penalty first, then solver based on data size |
fit_intercept / tol | Whether to fit intercept / convergence threshold | — | Usually keep defaults |
Experience with linear models: use a log grid (np.logspace(-4, 4, 20)) to scan C/alpha, and plot a "C vs cross-validation score" curve. If both ends of the curve are rising, the search range is insufficient — expand outward. If there's a plateau in the middle, the model is insensitive to regularization strength; take a conservative value in the middle of the plateau — no need to chase the exact optimum.
4. Neural Networks
Neural networks (MLP, CNN, Transformer) are the model family with the most severe hyperparameter coupling; see Deep Learning Fundamentals for mechanics. Their hyperparameters can be divided into three layers: capacity (how big), optimization (how to learn), and regularization (how to prevent overfitting).
| Hyperparameter | Meaning | Direction | Common Default | Notes |
|---|---|---|---|---|
learning_rate | Step size per update | Too large → diverges; too small → very slow convergence | Adam: 1e-3 / SGD: 1e-2 | Most important, tune first |
batch_size | Samples per batch | Larger = more stable gradient, more memory | 32–256 | Strongly coupled with learning rate, see Section V |
| Layer count & width per layer | Model capacity | Larger = stronger fit, more overfitting | Start with 2 layers, 64 wide | Start small, add more if insufficient |
dropout | Fraction of randomly dropped neurons | Larger = more overfit prevention | 0.2–0.5 | Start at 0.3 when overfitting |
weight_decay | L2 regularization strength on weights | Larger = more overfit prevention | 1e-6–1e-2 | Hidden regularization knob paired with Adam |
| Optimizer | Update rule (Adam/SGD+momentum) | — | Adam | Fix optimizer before tuning lr, don't keep jumping |
warmup_steps | Steps for learning rate warmup | Prevent early oscillation with large lr | 5%–10% of total steps | Essential with large batch + large lr |
scheduler | Strategy for learning rate decay during training | — | cosine / step decay | Key to convergence quality |
The "default value trap" in neural networks
The biggest mistake in neural network tuning is treating hyperparameters as independent variables to tune one by one. lr, batch_size, and weight_decay are strongly coupled — increasing lr by 10x alone might just push the model from "underfitting" to "diverging," which looks like an lr problem but is actually a failure to simultaneously adjust batch_size and weight_decay. This is why deep learning tuning is usually left to automated tools (see Optuna in Section IV), rather than manual trial-and-error.
5. Hyperparameter Panorama Summary Table
Lay out the "tune first, tune later" for the three model families side by side for easy comparison:
| Model Family | Tune First (high ROI) | Tune Next | Rarely Tune | Parameter Coupling |
|---|---|---|---|---|
| Tree/Ensemble | max_depth, min_child_weight, learning_rate | Sampling params, regularization | n_estimators (with early stopping) | Medium: lr couples with tree count |
| Linear | C / alpha (log grid) | l1_ratio, solver | Keep rest default | Low: almost independent |
| Neural Network | learning_rate, batch_size | Capacity, dropout, weight_decay | warmup/scheduler use defaults | High: lr ↔ batch ↔ wd all affect each other |
The selection logic behind this table: lower-coupling models suit manual + grid search; higher-coupling models suit Bayesian optimization. Next, we move to search methods.
III. Search Methods Comparison
1. Manual Tuning
Modify parameters one by one based on experience and intuition. Pros: controllable, builds domain intuition, nearly zero cost. Cons: depends on experience, hard to reproduce, easy to fall into local optima. Manual tuning is only truly viable when "you understand the model's hyperparameter terrain" — e.g., you know logistic regression's C uses a log scale, and tree model overfitting starts with min_child_weight.
2. Grid Search (Grid Search)
Given a set of candidate values for each hyperparameter, exhaustively compute their Cartesian product. scikit-learn's GridSearchCV is the standard implementation, with automatic cross-validation internally.
python
param_grid = {
"n_estimators": [100, 300],
"max_depth": [5, 10],
"min_samples_split": [2, 5],
} # 2 × 2 × 2 = 8 parameter combinations, × 5 folds = 40 full training runsPros: simple to implement, reproducible, naturally parallelizable. Cons: number of combinations explodes exponentially with parameter count (the curse of dimensionality) — 6 parameters with 10 candidates each means 1 million training runs, with most of the budget wasted in meaningless regions. It only suits scenarios with few parameters (≤2–3) where you have confidence in the value space.
3. Random Search
In their classic 2012 paper, Bergstra and Bengio proved a counterintuitive result: given the same budget, randomly sampling 60 combinations from a distribution almost always outperforms exhaustively enumerating a grid of the same size. The reason is simple: usually only two or three hyperparameters are truly decisive, and random sampling lets each parameter cover a wider range of values, whereas in grid search most combinations repeat across "unimportant dimensions."
python
from scipy.stats import randint, uniform
param_dist = {
"n_estimators": randint(50, 500), # uniform distribution over integers
"max_depth": randint(3, 20),
"max_features": uniform(0.1, 0.9), # uniform distribution over continuous values
}Pros: as simple to implement as grid search, immune to the curse of dimensionality, significantly more efficient. Cons: doesn't leverage historical experiment data; relies purely on probabilistic coverage. It's the default choice when "budget is limited + extreme optimality isn't required".
4. Bayesian Optimization
The core idea is using results from initial experiments to guide which parameters to try next. Specifically: a cheap surrogate model (usually Gaussian Process or Tree-structured Parzen Estimator, TPE) fits a black-box function "hyperparameters → validation score," then an acquisition function balances "exploitation (known good regions)" and "exploration (untried regions)" to select the next most promising candidate. Optuna, Hyperopt, and SMAC are all engineering implementations.
Bayesian optimization loop:
① Fit surrogate model from existing observations (params→scores)
② Acquisition function selects next candidate (balance exploit & explore)
③ Train and evaluate with candidate, get new observation
④ Go back to ① until budget exhaustedPros: highest efficiency for a given budget, especially when single training runs are expensive (deep learning, large ensembles). Cons: sequential exploration overhead (though Optuna supports parallel trials), higher barrier to implementation and understanding, more sensitive to parameter space definition.
5. Early Stopping & Iterative Methods
A method specific to deep learning and GBDT: instead of fully training every parameter combination, judge quality mid-training and terminate early. In deep learning, this means monitoring validation loss and stopping if it doesn't improve for N consecutive epochs; XGBoost/LightGBM use early_stopping_rounds for the same truncation on tree count. It's not a "search algorithm" per se, but rather an accelerator for all search methods — making each candidate ten times cheaper, enabling ten times more search iterations.
6. Summary Table and Selection Decision
| Method | Principle | Trials Needed | Pros | Cons | Use Case |
|---|---|---|---|---|---|
| Manual tuning | Experience + intuition | Minimal | Controllable, builds intuition | Depends on experience, hard to reproduce | Familiar model, very few params |
| Grid search | Exhaustive Cartesian product | Combinations × folds | Simple, parallelizable, reproducible | Curse of dimensionality, wastes budget | ≤2–3 params, discrete values |
| Random search | Random sampling from distributions | Comparable to grid | Broader coverage, simple to implement | Doesn't use historical info | 3–6 params, includes continuous params |
| Bayesian optimization | Surrogate model + acquisition function | Few (tens to hundreds) | Highest budget efficiency | Sequential overhead, complex implementation | Expensive training, many params, strong coupling |
| Early stopping | Truncate mid-training | Minimal | Lowest cost per candidate | Depends on dynamic eval protocol | Deep learning / GBDT |
Decision rules (select in this order):
- Single training run under a few seconds, params ≤3 → grid search;
- Params 3–6, medium budget → random search (always beats grid);
- Single training run minutes to hours, params ≥5, strong coupling → Bayesian optimization (Optuna);
- For any method, use early stopping when possible, multiplying your budget by ten in search iterations.
IV. Code Walkthroughs
Below, we run all three methods end-to-end on the same dataset (scikit-learn's built-in digits dataset) so the code can be copied and executed. Let's start with data preparation.
1. Data Preparation and Baseline
python
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# ── ① Data: built-in dataset, so data prep doesn't interfere with tuning ──
X, y = load_digits(return_X_y=True) # 1797 8×8 digit images
X_train, X_val, y_train, y_val = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print(f"Training: {X_train.shape}, Validation: {X_val.shape}")
# ── ② Baseline: use a "definitely runnable" set of default params first ──
baseline = RandomForestClassifier(n_estimators=100, random_state=42)
baseline.fit(X_train, y_train)
print(f"Baseline val accuracy: {accuracy_score(y_val, baseline.predict(X_val)):.4f}")Note that X_val is intentionally kept separate: cross-validation splits internally; X_val is for the final independent evaluation after all tuning is done — this is the key isolation preventing "tuning until you overfit the validation set," detailed in Section VI.
2. Grid Search: GridSearchCV
python
from sklearn.model_selection import GridSearchCV
param_grid = {
"n_estimators": [100, 300], # 2
"max_depth": [5, 10, None], # 3
"min_samples_split": [2, 5], # 2
} # 12 combos × 5 folds = 60 training runs
grid = GridSearchCV(
RandomForestClassifier(random_state=42, n_jobs=-1),
param_grid=param_grid,
cv=5, # 5-fold cross-validation
scoring="accuracy",
n_jobs=-1, # all parallel
verbose=1,
)
grid.fit(X_train, y_train)
print("Best params:", grid.best_params_)
print("CV best score: %.4f" % grid.best_score_)
print("Val score (independent eval): %.4f" % grid.score(X_val, y_val))Three outputs, each with its own meaning: best_params_ is the tuning conclusion; best_score_ is the average score in cross-validation (the decision basis for tuning); grid.score(X_val, y_val) is the final test on data that never participated in any decision.
3. Random Search: RandomizedSearchCV
python
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import randint, uniform
param_dist = {
"n_estimators": randint(50, 500),
"max_depth": randint(3, 20),
"min_samples_split": randint(2, 20),
"max_features": uniform(0.1, 1.0), # feature ratio, 0.1~1.0
}
random_search = RandomizedSearchCV(
RandomForestClassifier(random_state=42, n_jobs=-1),
param_distributions=param_dist,
n_iter=30, # only try 30 combos, not the full 450×17×18 space
cv=5,
scoring="accuracy",
n_jobs=-1,
random_state=42, # fixed random seed for reproducibility
)
random_search.fit(X_train, y_train)
print("Best params:", random_search.best_params_)
print("CV best score: %.4f" % random_search.best_score_)Comparing 30 combos × 5 folds = 150 training runs for random search: its "actual coverage" in high-dimensional space is far higher than the 12-combo grid search above — this is the direct engineering manifestation of the 2012 paper's conclusion.
4. Nested Cross-Validation: Preventing "Tuning Overfitting Itself"
The workflow above has a hidden risk: you selected hyperparameters AND estimated performance on X_train. If the tuning loop runs long enough, the best combination may simply be lucky — this is "tuning overfitting," and the model's score on X_val will be inflated. Rigorous evaluation uses nested cross-validation (nested CV): the outer fold estimates performance, the inner fold selects hyperparameters, and the two layers are strictly isolated. See Model Evaluation and Validation and Building a Model Evaluation Pipeline from Scratch for detailed theory.
python
from sklearn.model_selection import cross_val_score, KFold
# Outer 3 folds: each fold does a random search internally, then evaluates
outer_scores = []
outer_cv = KFold(n_splits=3, shuffle=True, random_state=42)
for train_idx, test_idx in outer_cv.split(X):
X_outer_train, X_outer_test = X[train_idx], X[test_idx]
y_outer_train, y_outer_test = y[train_idx], y[test_idx]
# Inner: search only on the outer training set
inner = RandomizedSearchCV(
RandomForestClassifier(random_state=42, n_jobs=-1),
param_distributions=param_dist,
n_iter=15, cv=3, scoring="accuracy", n_jobs=-1, random_state=42,
)
inner.fit(X_outer_train, y_outer_train)
outer_scores.append(accuracy_score(y_outer_test, inner.predict(X_outer_test)))
print("Nested CV scores: ", np.round(outer_scores, 4))
print("Unbiased performance estimate: %.4f (±%.4f)" % (np.mean(outer_scores), np.std(outer_scores)))The score from nested CV is an unbiased estimate of "if I use this tuning pipeline on new data, what score can I expect?" The cost is doubled computation, so run it once before final deployment, not during daily iterations.
5. Optuna Bayesian Optimization
When single training runs get expensive (deep learning, large models, large datasets), grid and random search are no longer cost-effective — use Optuna. It uses TPE (Tree-structured Parzen Estimator) as its surrogate model, supports pruning (i.e., early stopping truncation from Section III), parallel trials, and visualization.
bash
pip install optunapython
import optuna
from sklearn.model_selection import cross_val_score
def objective(trial):
# Declare each hyperparameter in Optuna's suggestion space
params = {
"n_estimators": trial.suggest_int("n_estimators", 100, 600, step=50),
"max_depth": trial.suggest_int("max_depth", 3, 18),
"min_samples_split": trial.suggest_int("min_samples_split", 2, 20),
"max_features": trial.suggest_float("max_features", 0.1, 1.0),
"criterion": trial.suggest_categorical("criterion", ["gini", "entropy"]),
}
clf = RandomForestClassifier(**params, random_state=42, n_jobs=-1)
# One trial = average score of a 5-fold cross-validation
return cross_val_score(clf, X_train, y_train, cv=5,
scoring="accuracy", n_jobs=-1).mean()
study = optuna.create_study(direction="maximize", study_name="rf_digits")
study.optimize(objective, n_trials=50) # 50 trials, learns automatically
print("Best hyperparams:", study.best_params)
print("Best score: %.4f" % study.best_value)Note that objective only uses X_train — as before, X_val is reserved for final evaluation. After running, you can use optuna.visualization.plot_param_importances(study) to see the importance ranking of each hyperparameter (based on permutation importance). This directly tells you "should you continue digging into a parameter next?" — this is Optuna's hidden bonus over brute-force search: it also tells you which knobs you no longer need to touch.
A note on data scale
All three code snippets above run in seconds on the 1797-sample digits dataset, making them easy to reproduce. In real projects, the more expensive a single training run is, the more you should focus your budget on "few and precise": first do a coarse scan with random search or manual tuning, lock in 2–3 sensitive parameters, then let Optuna dig deep on those few — don't let Optuna scan 8 parameters from the start.
V. Deep Learning Tuning
Deep learning's tuning philosophy is completely different from traditional machine learning: many parameters, strong coupling, expensive single training runs — you can almost never manually exhaust all combinations. Below are tuning strategies ranked by importance; see Optimization and Gradient Descent for optimization background and Deep Learning Fundamentals for network architecture background.
1. Learning Rate Always Comes First
In deep learning, learning rate (lr) has a far greater impact on final performance than any other hyperparameter: lr too large → loss oscillates and doesn't converge; lr too small → training stalls and you mistake it for insufficient capacity. Its optimal range typically spans 2–3 orders of magnitude, so it must be scanned on a log scale.
python
# Learning rate scan: train each lr for only 2~3 epochs, look at val loss trend
lr_candidates = [1e-4, 3e-4, 1e-3, 3e-3, 1e-2, 3e-2, 1e-1]
for lr in lr_candidates:
model = SimpleNet().to(device)
opt = torch.optim.Adam(model.parameters(), lr=lr)
for _ in range(3):
train_one_epoch(model, opt, train_loader)
val_loss = evaluate(model, val_loader)
print(f"lr={lr:6.0e} val_loss={val_loss:.4f}")Judgment criteria: validation loss steadily decreasing → lr is viable; oscillating and not dropping → lr too high; barely moving → lr too low. After a quick scan, select the best lr and move on.
Reference values for learning rates
- SGD + momentum: commonly 0.1 / 0.01;
- Adam: commonly 1e-3 / 3e-4;
- Large-scale pretraining / large batch: can reach 1e-2 or higher with warmup. These are just starting lines — scanning is always more reliable than memorizing values.
2. Warmup: Safety Belt for Early Training
If the learning rate jumps too high during training, it causes violent oscillation in the loss at the start, potentially pushing parameters to a point of no return. Warmup lets the learning rate climb linearly from near 0 to the target value before normal training begins — it's the standard companion for "large lr + large batch":
python
total_steps = n_epochs * len(train_loader)
warmup_steps = int(0.05 * total_steps) # warmup the first 5% of steps
def lr_lambda(step):
if step < warmup_steps: # linearly ramp to 1.0
return step / max(1, warmup_steps)
progress = (step - warmup_steps) / max(1, total_steps - warmup_steps)
return 0.5 * (1 + np.cos(np.pi * progress)) # then cosine anneal to 0
scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)3. Cosine Annealing: Let Learning Rate "Clock Out"
A constant lr alone isn't enough: in the later stages of training, a constant lr tends to wander near saddle points of the loss surface. Cosine annealing lets the learning rate smoothly drop to 0 following a cosine curve, allowing the model to make fine-grained convergence in the second half. The lr_lambda above already demonstrates the complete warmup + cosine combination — this is the de facto standard for pretraining and fine-tuning tasks (Loshchilov & Hutter, 2017).
Learning rate curve (one complete training run)
lr ┤
│ ╭───────────╮
│ ╱ ╲
│ ╱ cosine ╲
│ ╱──╱ ╲
│ ╱ warmup ╲
└──────────────────────────→ steps
0 ↑5% ↑100%4. The Relationship Between Batch Size and Learning Rate
This is the most practical rule in deep learning tuning: when batch size is multiplied by k, the learning rate should typically also be multiplied by k (the linear scaling rule), from Facebook AI's classic work Goyal et al., 2017:
Assume gradient is an approximation of true gradient plus random noise. Larger batch → less gradient noise → each step is more "accurate." To cover the same distance in the same number of steps, lr should scale up proportionally with batch. Standard practice: double batch → double lr; when doubling lr, also double warmup steps.
Two practical corollaries:
- Want to accelerate training with large batch? Scale up lr + warmup together, and you can typically halve training time with almost no accuracy loss;
- Intuitive comparison: small batch is "ask ten people a day" (noisy learning), large batch is "ask ten thousand people a day" (stable learning) — the latter is more expensive per step (slower), but the direction of each step is more accurate, so you can take larger strides (larger lr).
5. Remaining Key Points for Deep Learning Tuning
- dropout and weight_decay: when overfitting, tune weight_decay first (start at 1e-4) then dropout (start around 0.3); usually only one is needed;
- optimizer priority: use Adam first to quickly get reasonable results; if precision is insufficient, switch to SGD + momentum + cosine, which typically gains a few more points;
- change one coupled group at a time: lr, batch, and weight_decay should be treated as a unit — see Section VI;
- use automation when appropriate: in deep learning scenarios, changing
objectiveto "lr + batch_size + weight_decay + layer count" and letting Optuna run is an order of magnitude faster than manual trial-and-error.
VI. Tuning Discipline
No matter how advanced the methods are, if the experimental workflow isn't trustworthy, everything equals zero. Here are five iron rules; violating any of them can turn your "optimal parameters" into an illusion.
1. Change One Variable at a Time (or a Clearly Decoupled Small Group)
Changing two parameters and getting a good result means you don't know which one was responsible. While this often yields to automated search in modern workflows, this discipline still holds for manual iterations. If you must change two at once, explicitly write in your experiment log "this time changed A and B together because of reason C," and ensure the record is complete enough for post-hoc review.
2. Every Experiment Must Be Reproducible and Traceable
Tuning is essentially a "hypothesis-test" cycle; without records, there's no science. A minimum experiment record should at least include:
| Field | Description | Example |
|---|---|---|
| Experiment ID | Unique identifier | exp-2026-06-18-rf-depth |
| Parameter change | Difference from previous version | max_depth: 5 → 10 |
| Data and split | Data version, random seed | digits v1, seed=42 |
| Evaluation protocol | Fold count, metric | 5-fold, accuracy |
| Result | CV score ± standard deviation | 0.9712 ± 0.004 |
| Conclusion | Keep or discard this change? | Keep, proceed to next round |
A lazy but effective rule
Append all experiment changes and results to the same CSV/Notion table. After a week, the most valuable thing in your review isn't "which params were best," but "which paths were dead ends" — the latter saves you from walking the same path again next week.
3. Always Make Decisions on Validation, Never on Test
This is the most classic and most frequently violated discipline in machine learning. Data should be split into three parts:
Full dataset
├── Training set train → fit model
├── Validation set val → select hyperparameters (allowed repeated use)
└── Test set test → final one-time evaluation (use once and discard)Every hyperparameter adjustment is based on the validation set; the test set is used only once after all tuning is done. The reason is straightforward: as soon as you look at the test set more than twice, you start implicitly "memorizing" it — and the test set loses its meaning. For correct and incorrect evaluation protocol details, see Model Evaluation and Validation and Building a Model Evaluation Pipeline from Scratch.
4. Prevent the Tuning Process Itself from Overfitting the Validation Set
The tuning loop itself is a form of learning — it's "learning" the validation set. The longer the search runs and the more trials it makes, the more likely you are to select a combination that only got lucky on the validation set. Three lines of defense:
- Use random/Bayesian over exhaustive search as early as possible, reducing "gambling-style" sampling;
- Use nested cross-validation (demonstrated in Section IV) for key decisions to get unbiased estimates;
- Set a "distrust coefficient" for your final validation score: it's typically slightly higher than your nested CV average — don't take the small excess as truth.
5. Before Each Tuning Round, Ask "Is the ROI Worth It?"
The last and most important discipline: evaluate tuning by the clock, not by metrics. Spending 8 hours to go from 0.970 to 0.972 accuracy is nearly meaningless for an online business; the same 8 hours spent adding a neglected feature or fixing a data pipeline bug could bring a 5-point improvement. Tuning is the last-step fine-tuning for aligning a problem, not the problem itself. Data and features are the main event — this is the first principle of the "tuning mindset" from Section I that we want you to remember.
VII. Trade-offs and Considerations
Every choice in tuning is a trade-off between cost and benefit, worth making explicit:
- Compute budget vs. precision gain: tuning has clear diminishing returns. Plot a "search budget → validation score" curve, and you'll see it flatten quickly. The rational approach is to set a "good enough" threshold, stop when reached, and allocate remaining compute to other areas.
- Automation vs. interpretability: Optuna can automatically find good parameters, but you lose the understanding of "why this value is good." For small teams and long-term projects, it's worth keeping manual spot-checks alongside automated search — understanding the parameter terrain makes the next starting line much faster.
- Tuning vs. feature engineering: given the same budget, fixing a feature often yields an order of magnitude more gain than tuning a hyperparameter from 0.971 to 0.972. Feature engineering is at least as important as tuning — the site's core knowledge module has dedicated sections on it.
- Simple model + tuning vs. complex model + defaults: a well-tuned random forest often beats an XGBoost run with default params; a well-tuned 3-layer MLP beats an untuned ResNet 18. Matching hyperparameter effort to model complexity matters more than complexity itself.
- Search intensity vs. overfitting the validation set: the more you search, the more inflated your validation scores become. There's no free lunch for this contradiction; manage it only through nested CV and "use the test set once."
- Tool selection vs. team capability: whether to use Optuna or go manual depends on whether the team understands how to define the search space — an incorrect space definition (range too narrow, missing key parameters) makes any search method run in vain. See How to Choose Frameworks and Tools for the logic of tool and framework selection.
VIII. Further Reading
- Tuning presupposes understanding evaluation — Model Evaluation and Validation, Building a Model Evaluation Pipeline from Scratch
- Pitfalls to avoid in tuning (data leakage, overfitting validation set) — Common Pitfalls and Anti-Patterns
- The mechanism behind regularization hyperparameters — Overfitting and Regularization
- The principle behind tree model hyperparameters — Tree Models and Ensemble Learning
- The theoretical foundation for deep learning tuning — Optimization and Gradient Descent, Deep Learning Fundamentals
- The role of tuning within a complete project workflow — Building an ML Project from Scratch
- Terminology quick reference — Resource: Glossary
References
- scikit-learn: Tuning the hyper-parameters of an estimator (GridSearchCV / RandomizedSearchCV official docs)
- Optuna official documentation and tutorials
- Bergstra & Bengio. Random Search for Hyper-Parameter Optimization (JMLR 2012) — classic proof that "random search beats grid search"
- Snoek, Larochelle, Adams. Practical Bayesian Optimization of Machine Learning Algorithms (NeurIPS 2012) — pioneering work on Bayesian optimization for hyperparameter tuning
- Shahriari et al. Taking the Human Out of the Loop: A Review of Bayesian Optimization (Proceedings of the IEEE 2016) — authoritative review of Bayesian optimization
- Akiba et al. Optuna: A Next-generation Hyperparameter Optimization Framework (KDD 2019) — Optuna framework's original paper
- Loshchilov & Hutter. SGDR: Stochastic Gradient Descent with Warm Restarts (ICLR 2017) — cosine learning rate annealing
- Goyal et al. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (2017) — linear scaling rule for batch size and learning rate
- Smith. A Disciplined Approach to Neural Network Hyper-parameters: Part 1 – Learning Rate, Batch Size, Momentum, and Weight Decay (2018) — systematic experiments on coupling relationships between deep learning hyperparameters