Skip to content

Tuning Practice and Hyperparameter Optimization

Quick overview Hyperparameter tuning is an unavoidable step in deploying models, and also where compute and time are most easily wasted. This article explains the correct order for tuning, key hyperparameters for each model family, a comparison of grid/random/Bayesian search methods with code walkthroughs, and a complete methodology for deep learning tuning and experiment discipline.

Tuning Practice and Hyperparameter Optimization ​

Every minute you spend on tuning should first answer one question: is this parameter really worth tuning?

Hyperparameter optimization (HPO) is the most paradoxical part of machine learning engineering: it's described as critically important — "tuning is a craft," "the gap between experts and beginners comes down to this" — but the reality is that 90% of tuning in projects wastes compute, and what truly determines a model's ceiling is usually the data, features, and model choice. The goal of this article is not to teach you to memorize parameter values, but to help you build a complete methodology for "what to tune, when to tune it, what method to use, and how to ensure you're tuning correctly."

Let's start with a counterintuitive fact: a widely circulated Kaggle post found that top contestants spend far more time on feature engineering and data cleaning than on hyperparameter tuning — and when they do tune, they rarely use brute-force grid search, preferring instead "a few parameters, quick validation, narrowing by experience." This isn't because they've memorized parameter values, but because they put tuning in the right place.

The article is structured as follows: first, the tuning mindset (when you should actually tune); then a panoramic map of hyperparameters (what key knobs exist for each model family); then a comparison of four search methods; followed by complete code walkthroughs with sklearn and Optuna; then a deep dive into deep learning tuning; and finally, a set of tuning discipline rules to ensure your experiments are trustworthy.

I. Tuning Mindset: Data First, Then Model, Then Hyperparameters ​

1. Why "tuning from the start" is the biggest waste ​

The most common path for a beginner given a modeling task is: install XGBoost → run GridSearchCV overnight → stare at best_params_ with excitement → submit. The problem with this workflow isn't the tool, it's the order: hyperparameters are the last knob on the model, and four layers of problems lie above it that need to be investigated first.

┌─────────────────────────────────────────────────────────────┐
│  Problem-investigation order (top to bottom, higher = prior) │
├─────────────────────────────────────────────────────────────┤
│  ① Data layer    Enough samples? Labels correct? Leakage?     │
│     Normal distribution?                                      │
│  ② Feature layer Any useful features left unused? Missing/    │
│     outlier handling?                                         │
│  ③ Model layer   Right model family? Bias or variance issue?  │
│  ④ Hyperparam    Only after ①–③ are settled do hyperparams   │
│     get their turn                                            │
└─────────────────────────────────────────────────────────────┘

Why this order? Because higher-layer problems corrupt the signals at lower layers. If your data has leakage (e.g., deduplication before splitting), every "optimal value" you see at the hyperparameter level is fake — the model is just memorizing the answers, and you're painstakingly tuning the mouth shape for "memorizing the answers." For a complete list of fatal traps like data leakage, temporal leakage, and train/online inconsistency, see Common Pitfalls and Anti-Patterns; for a complete project workflow framework, see Building an ML Project from Scratch.

Don't tune on a shaky foundation

If you see high training metrics but poor validation performance — don't doubt your hyperparameters first, doubt overfitting and data leakage; if even the training metrics are low — don't rush to tune, first check whether your features and model match the problem. Hyperparameters can only amplify your gains once the foundation is solid; when the foundation is skewed, they only amplify errors.

2. Define "what tuning can solve" first ​

Hyperparameters can only affect one thing: the specific landing point of the bias-variance trade-off. Whether the model is underfitting (high bias) or overfitting (high variance) determines which direction you should turn:

SymptomDiagnosisTuning Direction
Training set & validation set both failUnderfitting / high biasIncrease capacity: deeper models/more trees, reduce regularization, lower min_samples_leaf
Training set perfect, validation set poorOverfitting / high varianceStrengthen regularization: add dropout/weight_decay, increase min_child_weight, reduce max_depth
Training set & validation set both near theoretical lower boundReached model ceilingStop tuning, go back to data and feature levels for gains

This judgment method is also called "bias-variance decomposition," the most practical engineering technique in Model Evaluation and Validation. Before any tuning, plot the training and validation errors against model capacity — you'll immediately know which segment of the curve you're on.

3. When hyperparameters actually get their turn ​

Check the following checklist before entering the tuning stage — only proceed when all conditions are met:

  1. A baseline model is running: a simple logistic regression or single decision tree already has reasonable performance, and you know what the "blind lower bound" is;
  2. Data and features are frozen: no more frequent adding/removing of features — otherwise, every previous tuning result is invalidated by a feature change;
  3. Evaluation protocol is fixed: splitting method, evaluation metrics, and cross-validation folds are all set in stone, see Building a Model Evaluation Pipeline from Scratch;
  4. Goal is clear: precision-prioritized (competitions, offline tasks) or latency/resource-prioritized (online inference). This determines the direction of the search space.

Tuning before meeting these conditions is like framing a painting on a wobbly easel.

II. Hyperparameter Panorama ​

1. Parameters vs. Hyperparameters ​

First, distinguish two concepts: parameters (parameters) are learned by the model from the data, such as coefficients w, b in linear regression or weights in neural networks — their values are determined by the training algorithm and don't need your intervention. Hyperparameters (hyperparameters) are the knobs you set before training, such as the number of trees, learning rate, and regularization strength — the data won't tell you automatically what values they should take, and that's exactly what tuning solves.

Parameters:  exist only after fit(X, y)    → determined by the optimization algorithm
Hyperparams: exist before fit(X, y)        → determined by the "tuner"

A practical rule: the fewer hyperparameters a model has, and the less sensitive it is to them, the easier it is to use; the stronger the interactions between hyperparameters, the more systematic search is needed. Random Forest has almost only one "any value works" parameter (n_estimators), which is why it's the universal baseline for Kaggle. Deep learning has seven or eight interdependent knobs, which is why it requires specialized tuning techniques.

2. Tree Models and Ensemble Methods ​

The tree model family (decision trees, random forests, XGBoost, LightGBM, CatBoost) is currently the de facto standard for tabular data; see Tree Models and Ensemble Learning for detailed mechanics. Their hyperparameters can be divided into four groups: structure (how deep each tree grows), ensemble (how many trees, how they sample), regularization (prevent overfitting), and training (learning rate, etc.).

HyperparameterMeaningDirectionCommon RangeTuning Tips
n_estimators / num_boost_roundNumber of treesMore = lower bias, but diminishing returns, linear time cost100–2000Tune last; early stopping is better than stacking trees
max_depthMaximum depth of a single treeDeeper = stronger fit, more overfitting3–15The #1 knob for controlling overfitting
learning_rate / etaStep size contributed by each treeSmaller = more stable, needs more trees0.01–0.3Coupled with tree count: halve lr, often need to double tree count
min_samples_split / min_child_weightMinimum samples/weight for splittingLarger = more overfit prevention2–20 / 1–10Prioritize increasing when overfitting
min_samples_leafMinimum samples per leafLarger = smoother trees1–20Closely related to small-sample datasets
max_featuresNumber of features considered per splitSmaller = lower variancesqrt or 0.3–0.7Effective to reduce when features are many
subsample / bagging_fractionSample ratio per treeSmaller = more overfit prevention0.6–1.0Common in GBDT; default for random forest
colsample_bytreeFeature ratio used per treeSame as above0.6–1.0XGBoost-specific, similar to max_features
lambda / reg_lambdaL2 regularization on leaf weightsLarger = more overfit prevention1–100Recommended when many trees, large depth

Tree model tuning order (time-saving golden sequence)

  1. Fix learning rate (default 0.1) and enough trees, tune max_depth and min_child_weight to a "underfit/overfit boundary";
  2. Then tune sampling params (subsample, colsample_bytree, max_features);
  3. Finally tune regularization (lambda) with early stopping to find optimal tree count. This order decouples interdependent parameters, moving only one or two knobs at each step.

3. Linear Models and Regularization ​

Linear models (linear regression, Ridge, Lasso, ElasticNet, logistic regression, linear SVM) have few but precise hyperparameters, almost all centered around "regularization strength". See Overfitting and Regularization for their mechanics and regularization mechanisms.

HyperparameterMeaningCommon RangeNotes
C (logistic regression, LinearSVC)Inverse of regularization strength: larger C = weaker penalty1e-4–1e4 (log grid)Easiest pitfall: C direction is opposite to "strength"
alpha (Ridge/Lasso/ElasticNet)Regularization strength: larger = stronger penalty1e-4–10Opposite direction to C, be careful
l1_ratio (ElasticNet)Mix ratio of L1 to L20–10 = pure Ridge, 1 = pure Lasso
penalty / solverRegularization type and solver algorithmSee sklearn docsUsually set penalty first, then solver based on data size
fit_intercept / tolWhether to fit intercept / convergence threshold—Usually keep defaults

Experience with linear models: use a log grid (np.logspace(-4, 4, 20)) to scan C/alpha, and plot a "C vs cross-validation score" curve. If both ends of the curve are rising, the search range is insufficient — expand outward. If there's a plateau in the middle, the model is insensitive to regularization strength; take a conservative value in the middle of the plateau — no need to chase the exact optimum.

4. Neural Networks ​

Neural networks (MLP, CNN, Transformer) are the model family with the most severe hyperparameter coupling; see Deep Learning Fundamentals for mechanics. Their hyperparameters can be divided into three layers: capacity (how big), optimization (how to learn), and regularization (how to prevent overfitting).

HyperparameterMeaningDirectionCommon DefaultNotes
learning_rateStep size per updateToo large → diverges; too small → very slow convergenceAdam: 1e-3 / SGD: 1e-2Most important, tune first
batch_sizeSamples per batchLarger = more stable gradient, more memory32–256Strongly coupled with learning rate, see Section V
Layer count & width per layerModel capacityLarger = stronger fit, more overfittingStart with 2 layers, 64 wideStart small, add more if insufficient
dropoutFraction of randomly dropped neuronsLarger = more overfit prevention0.2–0.5Start at 0.3 when overfitting
weight_decayL2 regularization strength on weightsLarger = more overfit prevention1e-6–1e-2Hidden regularization knob paired with Adam
OptimizerUpdate rule (Adam/SGD+momentum)—AdamFix optimizer before tuning lr, don't keep jumping
warmup_stepsSteps for learning rate warmupPrevent early oscillation with large lr5%–10% of total stepsEssential with large batch + large lr
schedulerStrategy for learning rate decay during training—cosine / step decayKey to convergence quality

The "default value trap" in neural networks

The biggest mistake in neural network tuning is treating hyperparameters as independent variables to tune one by one. lr, batch_size, and weight_decay are strongly coupled — increasing lr by 10x alone might just push the model from "underfitting" to "diverging," which looks like an lr problem but is actually a failure to simultaneously adjust batch_size and weight_decay. This is why deep learning tuning is usually left to automated tools (see Optuna in Section IV), rather than manual trial-and-error.

5. Hyperparameter Panorama Summary Table ​

Lay out the "tune first, tune later" for the three model families side by side for easy comparison:

Model FamilyTune First (high ROI)Tune NextRarely TuneParameter Coupling
Tree/Ensemblemax_depth, min_child_weight, learning_rateSampling params, regularizationn_estimators (with early stopping)Medium: lr couples with tree count
LinearC / alpha (log grid)l1_ratio, solverKeep rest defaultLow: almost independent
Neural Networklearning_rate, batch_sizeCapacity, dropout, weight_decaywarmup/scheduler use defaultsHigh: lr ↔ batch ↔ wd all affect each other

The selection logic behind this table: lower-coupling models suit manual + grid search; higher-coupling models suit Bayesian optimization. Next, we move to search methods.

III. Search Methods Comparison ​

1. Manual Tuning ​

Modify parameters one by one based on experience and intuition. Pros: controllable, builds domain intuition, nearly zero cost. Cons: depends on experience, hard to reproduce, easy to fall into local optima. Manual tuning is only truly viable when "you understand the model's hyperparameter terrain" — e.g., you know logistic regression's C uses a log scale, and tree model overfitting starts with min_child_weight.

Given a set of candidate values for each hyperparameter, exhaustively compute their Cartesian product. scikit-learn's GridSearchCV is the standard implementation, with automatic cross-validation internally.

python
param_grid = {
    "n_estimators": [100, 300],
    "max_depth": [5, 10],
    "min_samples_split": [2, 5],
}   # 2 × 2 × 2 = 8 parameter combinations, × 5 folds = 40 full training runs

Pros: simple to implement, reproducible, naturally parallelizable. Cons: number of combinations explodes exponentially with parameter count (the curse of dimensionality) — 6 parameters with 10 candidates each means 1 million training runs, with most of the budget wasted in meaningless regions. It only suits scenarios with few parameters (≤2–3) where you have confidence in the value space.

In their classic 2012 paper, Bergstra and Bengio proved a counterintuitive result: given the same budget, randomly sampling 60 combinations from a distribution almost always outperforms exhaustively enumerating a grid of the same size. The reason is simple: usually only two or three hyperparameters are truly decisive, and random sampling lets each parameter cover a wider range of values, whereas in grid search most combinations repeat across "unimportant dimensions."

python
from scipy.stats import randint, uniform
param_dist = {
    "n_estimators": randint(50, 500),      # uniform distribution over integers
    "max_depth": randint(3, 20),
    "max_features": uniform(0.1, 0.9),     # uniform distribution over continuous values
}

Pros: as simple to implement as grid search, immune to the curse of dimensionality, significantly more efficient. Cons: doesn't leverage historical experiment data; relies purely on probabilistic coverage. It's the default choice when "budget is limited + extreme optimality isn't required".

4. Bayesian Optimization ​

The core idea is using results from initial experiments to guide which parameters to try next. Specifically: a cheap surrogate model (usually Gaussian Process or Tree-structured Parzen Estimator, TPE) fits a black-box function "hyperparameters → validation score," then an acquisition function balances "exploitation (known good regions)" and "exploration (untried regions)" to select the next most promising candidate. Optuna, Hyperopt, and SMAC are all engineering implementations.

Bayesian optimization loop:
  ① Fit surrogate model from existing observations (params→scores)
  ② Acquisition function selects next candidate (balance exploit & explore)
  ③ Train and evaluate with candidate, get new observation
  ④ Go back to ① until budget exhausted

Pros: highest efficiency for a given budget, especially when single training runs are expensive (deep learning, large ensembles). Cons: sequential exploration overhead (though Optuna supports parallel trials), higher barrier to implementation and understanding, more sensitive to parameter space definition.

5. Early Stopping & Iterative Methods ​

A method specific to deep learning and GBDT: instead of fully training every parameter combination, judge quality mid-training and terminate early. In deep learning, this means monitoring validation loss and stopping if it doesn't improve for N consecutive epochs; XGBoost/LightGBM use early_stopping_rounds for the same truncation on tree count. It's not a "search algorithm" per se, but rather an accelerator for all search methods — making each candidate ten times cheaper, enabling ten times more search iterations.

6. Summary Table and Selection Decision ​

MethodPrincipleTrials NeededProsConsUse Case
Manual tuningExperience + intuitionMinimalControllable, builds intuitionDepends on experience, hard to reproduceFamiliar model, very few params
Grid searchExhaustive Cartesian productCombinations × foldsSimple, parallelizable, reproducibleCurse of dimensionality, wastes budget≤2–3 params, discrete values
Random searchRandom sampling from distributionsComparable to gridBroader coverage, simple to implementDoesn't use historical info3–6 params, includes continuous params
Bayesian optimizationSurrogate model + acquisition functionFew (tens to hundreds)Highest budget efficiencySequential overhead, complex implementationExpensive training, many params, strong coupling
Early stoppingTruncate mid-trainingMinimalLowest cost per candidateDepends on dynamic eval protocolDeep learning / GBDT

Decision rules (select in this order):

  1. Single training run under a few seconds, params ≤3 → grid search;
  2. Params 3–6, medium budget → random search (always beats grid);
  3. Single training run minutes to hours, params ≥5, strong coupling → Bayesian optimization (Optuna);
  4. For any method, use early stopping when possible, multiplying your budget by ten in search iterations.

IV. Code Walkthroughs ​

Below, we run all three methods end-to-end on the same dataset (scikit-learn's built-in digits dataset) so the code can be copied and executed. Let's start with data preparation.

1. Data Preparation and Baseline ​

python
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

# ── ① Data: built-in dataset, so data prep doesn't interfere with tuning ──
X, y = load_digits(return_X_y=True)          # 1797 8×8 digit images
X_train, X_val, y_train, y_val = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print(f"Training: {X_train.shape}, Validation: {X_val.shape}")

# ── ② Baseline: use a "definitely runnable" set of default params first ──
baseline = RandomForestClassifier(n_estimators=100, random_state=42)
baseline.fit(X_train, y_train)
print(f"Baseline val accuracy: {accuracy_score(y_val, baseline.predict(X_val)):.4f}")

Note that X_val is intentionally kept separate: cross-validation splits internally; X_val is for the final independent evaluation after all tuning is done — this is the key isolation preventing "tuning until you overfit the validation set," detailed in Section VI.

2. Grid Search: GridSearchCV ​

python
from sklearn.model_selection import GridSearchCV

param_grid = {
    "n_estimators": [100, 300],          # 2
    "max_depth": [5, 10, None],          # 3
    "min_samples_split": [2, 5],         # 2
}                                        # 12 combos × 5 folds = 60 training runs

grid = GridSearchCV(
    RandomForestClassifier(random_state=42, n_jobs=-1),
    param_grid=param_grid,
    cv=5,                                # 5-fold cross-validation
    scoring="accuracy",
    n_jobs=-1,                           # all parallel
    verbose=1,
)
grid.fit(X_train, y_train)

print("Best params:", grid.best_params_)
print("CV best score: %.4f" % grid.best_score_)
print("Val score (independent eval): %.4f" % grid.score(X_val, y_val))

Three outputs, each with its own meaning: best_params_ is the tuning conclusion; best_score_ is the average score in cross-validation (the decision basis for tuning); grid.score(X_val, y_val) is the final test on data that never participated in any decision.

3. Random Search: RandomizedSearchCV ​

python
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import randint, uniform

param_dist = {
    "n_estimators": randint(50, 500),
    "max_depth": randint(3, 20),
    "min_samples_split": randint(2, 20),
    "max_features": uniform(0.1, 1.0),   # feature ratio, 0.1~1.0
}

random_search = RandomizedSearchCV(
    RandomForestClassifier(random_state=42, n_jobs=-1),
    param_distributions=param_dist,
    n_iter=30,               # only try 30 combos, not the full 450×17×18 space
    cv=5,
    scoring="accuracy",
    n_jobs=-1,
    random_state=42,         # fixed random seed for reproducibility
)
random_search.fit(X_train, y_train)

print("Best params:", random_search.best_params_)
print("CV best score: %.4f" % random_search.best_score_)

Comparing 30 combos × 5 folds = 150 training runs for random search: its "actual coverage" in high-dimensional space is far higher than the 12-combo grid search above — this is the direct engineering manifestation of the 2012 paper's conclusion.

4. Nested Cross-Validation: Preventing "Tuning Overfitting Itself" ​

The workflow above has a hidden risk: you selected hyperparameters AND estimated performance on X_train. If the tuning loop runs long enough, the best combination may simply be lucky — this is "tuning overfitting," and the model's score on X_val will be inflated. Rigorous evaluation uses nested cross-validation (nested CV): the outer fold estimates performance, the inner fold selects hyperparameters, and the two layers are strictly isolated. See Model Evaluation and Validation and Building a Model Evaluation Pipeline from Scratch for detailed theory.

python
from sklearn.model_selection import cross_val_score, KFold

# Outer 3 folds: each fold does a random search internally, then evaluates
outer_scores = []
outer_cv = KFold(n_splits=3, shuffle=True, random_state=42)

for train_idx, test_idx in outer_cv.split(X):
    X_outer_train, X_outer_test = X[train_idx], X[test_idx]
    y_outer_train, y_outer_test = y[train_idx], y[test_idx]

    # Inner: search only on the outer training set
    inner = RandomizedSearchCV(
        RandomForestClassifier(random_state=42, n_jobs=-1),
        param_distributions=param_dist,
        n_iter=15, cv=3, scoring="accuracy", n_jobs=-1, random_state=42,
    )
    inner.fit(X_outer_train, y_outer_train)
    outer_scores.append(accuracy_score(y_outer_test, inner.predict(X_outer_test)))

print("Nested CV scores: ", np.round(outer_scores, 4))
print("Unbiased performance estimate: %.4f (±%.4f)" % (np.mean(outer_scores), np.std(outer_scores)))

The score from nested CV is an unbiased estimate of "if I use this tuning pipeline on new data, what score can I expect?" The cost is doubled computation, so run it once before final deployment, not during daily iterations.

5. Optuna Bayesian Optimization ​

When single training runs get expensive (deep learning, large models, large datasets), grid and random search are no longer cost-effective — use Optuna. It uses TPE (Tree-structured Parzen Estimator) as its surrogate model, supports pruning (i.e., early stopping truncation from Section III), parallel trials, and visualization.

bash
pip install optuna
python
import optuna
from sklearn.model_selection import cross_val_score

def objective(trial):
    # Declare each hyperparameter in Optuna's suggestion space
    params = {
        "n_estimators": trial.suggest_int("n_estimators", 100, 600, step=50),
        "max_depth": trial.suggest_int("max_depth", 3, 18),
        "min_samples_split": trial.suggest_int("min_samples_split", 2, 20),
        "max_features": trial.suggest_float("max_features", 0.1, 1.0),
        "criterion": trial.suggest_categorical("criterion", ["gini", "entropy"]),
    }
    clf = RandomForestClassifier(**params, random_state=42, n_jobs=-1)
    # One trial = average score of a 5-fold cross-validation
    return cross_val_score(clf, X_train, y_train, cv=5,
                           scoring="accuracy", n_jobs=-1).mean()

study = optuna.create_study(direction="maximize", study_name="rf_digits")
study.optimize(objective, n_trials=50)      # 50 trials, learns automatically

print("Best hyperparams:", study.best_params)
print("Best score: %.4f" % study.best_value)

Note that objective only uses X_train — as before, X_val is reserved for final evaluation. After running, you can use optuna.visualization.plot_param_importances(study) to see the importance ranking of each hyperparameter (based on permutation importance). This directly tells you "should you continue digging into a parameter next?" — this is Optuna's hidden bonus over brute-force search: it also tells you which knobs you no longer need to touch.

A note on data scale

All three code snippets above run in seconds on the 1797-sample digits dataset, making them easy to reproduce. In real projects, the more expensive a single training run is, the more you should focus your budget on "few and precise": first do a coarse scan with random search or manual tuning, lock in 2–3 sensitive parameters, then let Optuna dig deep on those few — don't let Optuna scan 8 parameters from the start.

V. Deep Learning Tuning ​

Deep learning's tuning philosophy is completely different from traditional machine learning: many parameters, strong coupling, expensive single training runs — you can almost never manually exhaust all combinations. Below are tuning strategies ranked by importance; see Optimization and Gradient Descent for optimization background and Deep Learning Fundamentals for network architecture background.

1. Learning Rate Always Comes First ​

In deep learning, learning rate (lr) has a far greater impact on final performance than any other hyperparameter: lr too large → loss oscillates and doesn't converge; lr too small → training stalls and you mistake it for insufficient capacity. Its optimal range typically spans 2–3 orders of magnitude, so it must be scanned on a log scale.

python
# Learning rate scan: train each lr for only 2~3 epochs, look at val loss trend
lr_candidates = [1e-4, 3e-4, 1e-3, 3e-3, 1e-2, 3e-2, 1e-1]
for lr in lr_candidates:
    model = SimpleNet().to(device)
    opt = torch.optim.Adam(model.parameters(), lr=lr)
    for _ in range(3):
        train_one_epoch(model, opt, train_loader)
    val_loss = evaluate(model, val_loader)
    print(f"lr={lr:6.0e}  val_loss={val_loss:.4f}")

Judgment criteria: validation loss steadily decreasing → lr is viable; oscillating and not dropping → lr too high; barely moving → lr too low. After a quick scan, select the best lr and move on.

Reference values for learning rates

  • SGD + momentum: commonly 0.1 / 0.01;
  • Adam: commonly 1e-3 / 3e-4;
  • Large-scale pretraining / large batch: can reach 1e-2 or higher with warmup. These are just starting lines — scanning is always more reliable than memorizing values.

2. Warmup: Safety Belt for Early Training ​

If the learning rate jumps too high during training, it causes violent oscillation in the loss at the start, potentially pushing parameters to a point of no return. Warmup lets the learning rate climb linearly from near 0 to the target value before normal training begins — it's the standard companion for "large lr + large batch":

python
total_steps = n_epochs * len(train_loader)
warmup_steps = int(0.05 * total_steps)        # warmup the first 5% of steps

def lr_lambda(step):
    if step < warmup_steps:                    # linearly ramp to 1.0
        return step / max(1, warmup_steps)
    progress = (step - warmup_steps) / max(1, total_steps - warmup_steps)
    return 0.5 * (1 + np.cos(np.pi * progress))  # then cosine anneal to 0

scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)

3. Cosine Annealing: Let Learning Rate "Clock Out" ​

A constant lr alone isn't enough: in the later stages of training, a constant lr tends to wander near saddle points of the loss surface. Cosine annealing lets the learning rate smoothly drop to 0 following a cosine curve, allowing the model to make fine-grained convergence in the second half. The lr_lambda above already demonstrates the complete warmup + cosine combination — this is the de facto standard for pretraining and fine-tuning tasks (Loshchilov & Hutter, 2017).

Learning rate curve (one complete training run)
lr ┤
   │        ╭───────────╮
   │       ╱             ╲
   │      ╱   cosine      ╲
   │  ╱──╱                 ╲
   │ ╱  warmup               ╲
   └──────────────────────────→ steps
   0   ↑5%               ↑100%

4. The Relationship Between Batch Size and Learning Rate ​

This is the most practical rule in deep learning tuning: when batch size is multiplied by k, the learning rate should typically also be multiplied by k (the linear scaling rule), from Facebook AI's classic work Goyal et al., 2017:

Assume gradient is an approximation of true gradient plus random noise. Larger batch → less gradient noise → each step is more "accurate." To cover the same distance in the same number of steps, lr should scale up proportionally with batch. Standard practice: double batch → double lr; when doubling lr, also double warmup steps.

Two practical corollaries:

  • Want to accelerate training with large batch? Scale up lr + warmup together, and you can typically halve training time with almost no accuracy loss;
  • Intuitive comparison: small batch is "ask ten people a day" (noisy learning), large batch is "ask ten thousand people a day" (stable learning) — the latter is more expensive per step (slower), but the direction of each step is more accurate, so you can take larger strides (larger lr).

5. Remaining Key Points for Deep Learning Tuning ​

  • dropout and weight_decay: when overfitting, tune weight_decay first (start at 1e-4) then dropout (start around 0.3); usually only one is needed;
  • optimizer priority: use Adam first to quickly get reasonable results; if precision is insufficient, switch to SGD + momentum + cosine, which typically gains a few more points;
  • change one coupled group at a time: lr, batch, and weight_decay should be treated as a unit — see Section VI;
  • use automation when appropriate: in deep learning scenarios, changing objective to "lr + batch_size + weight_decay + layer count" and letting Optuna run is an order of magnitude faster than manual trial-and-error.

VI. Tuning Discipline ​

No matter how advanced the methods are, if the experimental workflow isn't trustworthy, everything equals zero. Here are five iron rules; violating any of them can turn your "optimal parameters" into an illusion.

1. Change One Variable at a Time (or a Clearly Decoupled Small Group) ​

Changing two parameters and getting a good result means you don't know which one was responsible. While this often yields to automated search in modern workflows, this discipline still holds for manual iterations. If you must change two at once, explicitly write in your experiment log "this time changed A and B together because of reason C," and ensure the record is complete enough for post-hoc review.

2. Every Experiment Must Be Reproducible and Traceable ​

Tuning is essentially a "hypothesis-test" cycle; without records, there's no science. A minimum experiment record should at least include:

FieldDescriptionExample
Experiment IDUnique identifierexp-2026-06-18-rf-depth
Parameter changeDifference from previous versionmax_depth: 5 → 10
Data and splitData version, random seeddigits v1, seed=42
Evaluation protocolFold count, metric5-fold, accuracy
ResultCV score ± standard deviation0.9712 ± 0.004
ConclusionKeep or discard this change?Keep, proceed to next round

A lazy but effective rule

Append all experiment changes and results to the same CSV/Notion table. After a week, the most valuable thing in your review isn't "which params were best," but "which paths were dead ends" — the latter saves you from walking the same path again next week.

3. Always Make Decisions on Validation, Never on Test ​

This is the most classic and most frequently violated discipline in machine learning. Data should be split into three parts:

Full dataset
├── Training set train     → fit model
├── Validation set val     → select hyperparameters (allowed repeated use)
└── Test set test          → final one-time evaluation (use once and discard)

Every hyperparameter adjustment is based on the validation set; the test set is used only once after all tuning is done. The reason is straightforward: as soon as you look at the test set more than twice, you start implicitly "memorizing" it — and the test set loses its meaning. For correct and incorrect evaluation protocol details, see Model Evaluation and Validation and Building a Model Evaluation Pipeline from Scratch.

4. Prevent the Tuning Process Itself from Overfitting the Validation Set ​

The tuning loop itself is a form of learning — it's "learning" the validation set. The longer the search runs and the more trials it makes, the more likely you are to select a combination that only got lucky on the validation set. Three lines of defense:

  • Use random/Bayesian over exhaustive search as early as possible, reducing "gambling-style" sampling;
  • Use nested cross-validation (demonstrated in Section IV) for key decisions to get unbiased estimates;
  • Set a "distrust coefficient" for your final validation score: it's typically slightly higher than your nested CV average — don't take the small excess as truth.

5. Before Each Tuning Round, Ask "Is the ROI Worth It?" ​

The last and most important discipline: evaluate tuning by the clock, not by metrics. Spending 8 hours to go from 0.970 to 0.972 accuracy is nearly meaningless for an online business; the same 8 hours spent adding a neglected feature or fixing a data pipeline bug could bring a 5-point improvement. Tuning is the last-step fine-tuning for aligning a problem, not the problem itself. Data and features are the main event — this is the first principle of the "tuning mindset" from Section I that we want you to remember.

VII. Trade-offs and Considerations ​

Every choice in tuning is a trade-off between cost and benefit, worth making explicit:

  • Compute budget vs. precision gain: tuning has clear diminishing returns. Plot a "search budget → validation score" curve, and you'll see it flatten quickly. The rational approach is to set a "good enough" threshold, stop when reached, and allocate remaining compute to other areas.
  • Automation vs. interpretability: Optuna can automatically find good parameters, but you lose the understanding of "why this value is good." For small teams and long-term projects, it's worth keeping manual spot-checks alongside automated search — understanding the parameter terrain makes the next starting line much faster.
  • Tuning vs. feature engineering: given the same budget, fixing a feature often yields an order of magnitude more gain than tuning a hyperparameter from 0.971 to 0.972. Feature engineering is at least as important as tuning — the site's core knowledge module has dedicated sections on it.
  • Simple model + tuning vs. complex model + defaults: a well-tuned random forest often beats an XGBoost run with default params; a well-tuned 3-layer MLP beats an untuned ResNet 18. Matching hyperparameter effort to model complexity matters more than complexity itself.
  • Search intensity vs. overfitting the validation set: the more you search, the more inflated your validation scores become. There's no free lunch for this contradiction; manage it only through nested CV and "use the test set once."
  • Tool selection vs. team capability: whether to use Optuna or go manual depends on whether the team understands how to define the search space — an incorrect space definition (range too narrow, missing key parameters) makes any search method run in vain. See How to Choose Frameworks and Tools for the logic of tool and framework selection.

VIII. Further Reading ​

References ​