Skip to content

DL Design Principles

Quick overview Deep learning engineering isn't guesswork — it's a trainable set of mental habits. This article distills eight design principles: start with a simple baseline, data first, modularity and configurability, change one variable at a time, reproducibility, progressive complexity, default to suspecting your own code, and scalability thinking — with concrete actions for each.

DL Design Principles ​

One-sentence definition: the design principles of deep learning are the habit of institutionalizing the "change what, verify how, attribute to what" methodology, so that the ROI of every experiment is maximized — they matter more for your ceiling than any single model trick.

These principles are scattered across the practice articles; this article distills them into eight actionable principles and weaves them together.

Principle 1: Get a Simple Baseline Running First, Then Upgrade ​

Principle: for any task, first build a "definitely runs, mediocre results" baseline, then upgrade item by item.

The simplest, most reliable baselines include: zero-rule (always predict the majority class), linear models (logistic regression), shallow MLPs, zero-shot or direct inference from official pre-trained models. The value of baselines:

  • Provide a "reference frame" for all subsequent improvements: if an enhancement only beats the baseline by 1%, the problem isn't with the enhancement.
  • Verify pipeline correctness: getting the baseline to run = data, training, and evaluation chains are all working.
  • Quickly rule out "the task is inherently unsolvable": if the simplest model reaches 90% on a subset, the task is learnable.

Baselines are the first puzzle piece

The very first step of building a project from scratch should be the baseline — the v1 of progressive tutorials is a textbook example.

Principle 2: Data First — Look at the Data, Then Pick the Model ​

Principle: before writing any model code, spend time looking at the data and understanding its "personality."

Concrete actions:

  1. Visualize samples: for classification, draw 20–50 images/text snippets and manually check labels against content.
  2. Check distributions: class balance, value ranges, missing values, outliers, duplicates.
  3. Think "how would a human solve this?": what clues do humans rely on for this task? Can the model see those clues (resolution, context length, feature engineering)?
  4. Define evaluation: before choosing a model, clearly define "what counts as good" (see Evaluation Practices).

A typical counterexample: the model stubbornly fails to learn a certain class, you spend a week tinkering with the architecture, only to discover that half the labels for that class are wrong. Architecture can never compensate for data problems. For systematic data engineering methods, see Data and Data Engineering.

Principle 3: Modular and Configurable Experiments ​

Principle: make experiments "parameterized" so that each experiment is a config change, not a code rewrite.

Two layers:

  1. Code modularity: data, model, training, evaluation, and visualization each live in independent modules (structure template in Build Your Own DL Project). Change one thing without affecting the rest.
  2. Hyperparameter configurability: all hyperparameters go into config.yaml or CLI arguments, no hardcoding in code:
yaml
# config.yaml
data:
  dataset: cifar10
  batch_size: 128
  augmentation: [random_crop, hflip]
model:
  arch: resnet18
  pretrained: true
train:
  optimizer: adam
  lr: 1e-3
  scheduler: cosine_with_warmup
  epochs: 30
  seed: 42
python
import yaml, argparse
cfg = yaml.safe_load(open("config.yaml"))
# ... read with cfg["train"]["lr"], etc.

Configurability brings three direct benefits: experiments are reproducible (change one line of yaml and you have a new experiment), trackable (log the full config), and automatable (scripts batch-replace configs for grid/Bayesian search).

Principle 4: Experimental Discipline — Change One Variable at a Time ​

Principle: change only one factor per experiment, otherwise results are unattributable.

This sounds simple but is extremely hard in practice — because the temptation is strong: "while I'm at it, I'll also tweak the learning rate." Result: the model improves and you don't know why, or it gets worse and you don't know which change caused it.

Correct workflow:

  1. Record the baseline before changing (metrics, config, seed).
  2. Change only one variable, lock everything else (including random seeds, unless you're specifically studying seed effects).
  3. Record results after running, compare to the baseline.
  4. Repeat, rather than changing multiple things in parallel.

Variable lock checklist

Random seeds, data versions, code commits, framework versions, GPU models — all of these are "variables." Changing one at a time means freezing all of them too. For the full reproducibility checklist, see Training Recipes and Hyperparameter Tuning.

Principle 5: Reproducibility — Seeds, Versions, Documentation ​

Principle: the output of an experiment is not "model weights," it's "evidence that it can be reproduced."

  • Fix seeds: torch.manual_seed + numpy + random + cudnn.deterministic (code in Training Recipes and Hyperparameter Tuning).
  • Fix versions: pin requirements.txt (exact versions), record Python/CUDA/PyTorch versions.
  • Document: README should spell out installation, reproduction commands, data sources, hyperparameter configs; experiment results (curves, metric tables) should be saved to disk.
  • Golden reproduction: before shipping or submitting, re-run from scratch with the same seed and confirm the numbers match.

The value of reproducibility shows up within two weeks: you'll forget how you tuned things, and the README + config will pull you back in.

Principle 6: Progressive Complexity — MLP → CNN → Pre-trained ​

Principle: model complexity should scale with "task evidence," not jump to the fanciest config from the start.

Standard ramp-up path:

Logistic regression/MLP → Small CNN/shallow Transformer → Deep architecture → Pre-trained transfer → Larger scale/search
   Verify pipeline          Verify inductive bias              Verify scale effects     Leverage external knowledge     Diminishing returns zone

The criterion for staying at each level: has the validation metric at the current level clearly hit a bottleneck (training loss is already low, validation stops improving, errors concentrate on a specific class)? The rationale for level jumps should come from evidence, not from "everyone's using ResNet."

This principle was demonstrated in Progressive Tutorial: Three Versions: MLP 40% → CNN 78% → pre-trained 90%+. It avoids both extremes: over-engineering (using a 100-layer model for a task with 5,000 images) and premature abandonment (concluding a task is unsolvable just because MLP performs poorly).

Principle 7: When Things Fail, Suspect Your Own Code, Not the Model ​

Principle: when results don't match expectations, default to assuming "my code/data has a bug" rather than "the model is bad" or "the task is too hard."

Why this counterintuitive principle is so effective:

  • Model design flaws typically manifest as "explainably bad" (you can analyze why), while code bugs typically manifest as "mysteriously bad."
  • Debugging cost asymmetry: 30 minutes to check code vs. a week to rewrite the model under the false assumption that the architecture is the problem.
  • Deep learning frameworks are too "permissive": shape mismatches will throw errors, but semantic bugs (misaligned labels, wrong masks, incorrect normalization stats) silently pass with completely wrong results.

Practical approach: first use the "overfit to 100% on a small subset" technique from Debugging and Diagnosis to rule out pipeline issues, then do gradient checking, ablation studies, and only then suspect the model design itself. When the model performs so poorly that it "makes no sense," it's a code problem 90% of the time (full anti-pattern list in Common Pitfalls and Anti-patterns).

Principle 8: Scalability Thinking — Every Line You Write Might Be Reused ​

Principle: default to writing code as if "someone else (including you in three months) will run, modify, and extend this."

Concrete habits:

  • Functions over scripts: turn training loops, evaluation, and visualization into reusable functions, not scattered top-level code.
  • Device-agnostic: read device from config so the code runs on both CPU and GPU.
  • Explicit over implicit: spell out shapes, names, and comments; extract magic numbers into constants.
  • Design for "the next dataset": decouple the dataset interface in data.py from concrete datasets; swapping datasets only requires a config change.
  • Extract shared layers early: abstract common training/evaluation skeletons for multi-task use (but avoid over-abstraction — the previous principles can clash with this one).

Scalability isn't about writing a bunch of unused abstractions for the future; it's about keeping iteration cost growing approximately linearly as the project grows — which is exactly the core standard for evaluating code quality in Portfolio Projects.

Trade-offs and Boundaries: Principles Will Clash ​

These eight principles aren't blindly followed in full:

  • Progressive complexity vs. time budget: when competitions or deliverables have a deadline, "jump straight to pre-trained + known recipe" is often the better strategy, keeping the baseline in mind rather than running it out.
  • Change one variable at a time vs. experiment budget: true hyperparameter search (grid/Bayesian) is essentially "multi-variable in parallel," but attribution is guaranteed by the search method, not single-variable discipline.
  • Scalability vs. rapid iteration: during research exploration, simpler code is always better; over-engineering makes every modification slower.

There is only one criterion: the goal at the current stage is "getting trustworthy answers fastest." Exploration phases lean toward simple and direct; delivery phases lean toward engineering rigor. At the end of each phase, ask yourself: if someone else (or you in two months) picks this up, is this code and documentation sufficient?

Further Reading ​

References ​