Skip to content

ML Design Principles

Quick overview ML project failures rarely stem from models not being advanced enough; more often they die from design errors. This article distills ten design principles: run baselines first, simplicity first, evaluation first, data > model, reproducibility, prevent leakage, monitor at deployment, error-analysis-driven iteration, with exceptions and team norms discussed.

ML Design Principles ​

One-sentence version: the success or failure of an ML project is decided before a single line of model code is written.

Beginners often focus on "which model to use" — XGBoost or Transformer? Tuning or switching architecture? But real project post-mortems repeatedly tell us: models are rarely the bottleneck; design decisions are. Wrong data splitting method, and all subsequent experiments are wasted; no baseline, and you have no idea whether "improvement" is real or not; no evaluation protocol, and two models can't be compared; deploying without monitoring, and the model dies unnoticed for three months. Google's Rules of ML condenses this idea into one sentence:

"The most common cause of ML project failure is not algorithmic problems, but system design problems."

This article gives ten actionable design principles, each answering three questions: why this principle matters, how to do it right, and what the most common counterexample looks like. They don't require you to have learned all algorithms first — quite the opposite, most of these principles are written before algorithms.

Positioning of this article

This is a "process-level" article: it talks about how to organize an ML project, not how to train a specific model. Read Building an ML Project from Scratch first to build the overall workflow, then read this article to add "discipline" to each step in the workflow; the negative list counterpart is in Common Pitfalls and Anti-Patterns.

I. Ten Principles at a Glance ​

#PrincipleOne-LinerCounterexample Signal
1Run baselines firstBefore optimizing, have something to compare toJump straight into complex models with no reference point
2Simplicity firstStart simple; complexity needs evidenceBlindly using deep learning, can't run or explain
3Evaluation firstDefine evaluation protocol before writing any modeling codeFinished training the model before thinking "how do we score?"
4Data > modelData quality and volume impact results more than model choiceThree months of tuning but no week spent cleaning data
5Reproducibility is non-negotiableSame code and data must reproduce the same resultMetrics change on a different machine, nobody knows why
6Always prevent leakageNever use at training time what won't be available at inference99-point offline metrics, immediately crashes on deployment
7Deployment must include monitoringModel deployment is the beginning, not the endDeployed and ignored, metrics quietly collapse three months later
8Iteration driven by error analysisLet evidence from failing samples drive improvement, not gut feelingsRandomly switching models without looking at what's wrong
9Write docs and commentsThe you three months from now is a stranger tooOnly experiment outputs are model_final_v3_really_final.ipynb
10Understand business goalsOffline metrics ≠ business valueAUC went up 0.02 offline, business sees zero change

These ten aren't isolated; they form an iteration loop:

        ┌───────────────────────────────────────────┐
        │      Business goals (Principle 10)         │
        └───────────────────┬───────────────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │  Evaluation protocol first (Principle 3)   │
        │  Split method / metric definition / baseline│
        └───────────────────┬───────────────────────┘
                            ▼
   ┌─────────────┬──────────┴───────────┬───────────────┐
   ▼             ▼                      ▼               ▼
Data health    Simple baseline      Feature & data    Error analysis
(Principle 4)  (Principles 1, 2)    Engineering       (Principle 8)
                                 (Principles 4, 6)
   └─────────────┴──────────┬───────────┴───────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │  Reproducible experiment logs              │
        │  (Principles 5, 9)                        │
        └───────────────────┬───────────────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │  Deploy + monitor (Principle 7) → back to error analysis │
        └───────────────────────────────────────────┘

A common misunderstanding

"Design principles" sounds like theoretical preaching, but it's actually a time-saving methodology. Skipping baselines, you might spend two weeks on an "improvement," only to find it doesn't even beat the baseline; skipping evaluation design, your experiment conclusions can't withstand questioning. Principles aren't constraints; they're anchors pulling you out of "seemingly progressing, actually spinning wheels."

II. Detailed Breakdown ​

Principle 1: Establish a Baseline Before Talking About Optimization ​

Why. Without a baseline, "improvement" has no reference frame. You think a new model improved accuracy from 83% to 85%, but if you never ran a "learn nothing" baseline, you have no idea whether 83% itself is good or bad — maybe 90% of samples belong to the majority class, and no model can beat "always predict majority." Baselines also give you a cost anchor: if the simplest method already meets business requirements, all subsequent complex work is superfluous. Google's Rules of ML gives its #1 rule as: "Before any optimization, get a simple model live (or at least running) and measure its performance."

How to do it. There are three levels of baselines, escalating:

  1. Dummy baseline: majority-class classifier, mean regressor, random prediction. A one-liner with scikit-learn's DummyClassifier.
  2. Simple model baseline: logistic regression, linear regression, single decision tree, all with default parameters.
  3. Minimal domain baseline: if the problem has industry precedents (e.g., click-through rate prediction using LR), use that precedent as the base.
python
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

X, y = ...  # your features and labels

# Dummy baseline: always predict majority class
dummy = DummyClassifier(strategy="most_frequent")
dummy_score = cross_val_score(dummy, X, y, cv=5).mean()

# Simple baseline: default-parameter logistic regression
lr = LogisticRegression(max_iter=1000)
lr_score = cross_val_score(lr, X, y, cv=5).mean()

print(f"Dummy baseline : {dummy_score:.4f}")
print(f"Logistic regres: {lr_score:.4f}")
# Compare any subsequent model against these two numbers first

Counterexample. The most common counterexample is "skipping the baseline and jumping straight to the strongest model": only 50k rows of data, but deploying a hundreds-layer Chinese BERT fine-tune, training for a week, 200ms inference, unexplainable — only to find that logistic regression was only 0.3 percentage points behind under the same evaluation protocol, at a thousand times lower cost. The second counterexample is "baseline not recorded": ran the baseline but didn't leave numbers and code version; two weeks later, can't compare or reproduce, which equals not having run it.

Principle 2: Simple Models First (Occam's Razor in ML) ​

Why. Occam's razor says "do not multiply entities beyond necessity"; the ML translation is: when effects are similar, always choose the simpler one. Simple models have four systemic advantages:

  • Interpretable: linear model coefficients, decision tree paths — can be directly explained to business stakeholders, quickly traced when things go wrong.
  • Stable: complex models are more sensitive to distribution drift, noise, and implementation details; simple models are "tough."
  • Cheap: fast to train, fast to infer, low ops cost, iteration speed is multiples of complex models.
  • Easier to diagnose: when simple models fail, you can tell if it's a data problem or feature problem; with deep models, you can only guess.

How to do it. Move up the "complexity ladder" step by step, upgrading only with evidence beyond the current level:

Logistic/linear regression → Regularized LR → Single decision tree → RF/GBDT → Shallow NN → Deep learning
    ●                                                                    ●
  Start: always begin here                              End: don't reach here without evidence

Evidence for upgrading includes: simple models have converged to a bottleneck (error analysis shows insufficient expressive power), train/val gap indicates underfitting and more data provides no help. In practice, tabular data often plateaus at gradient-boosted trees (XGBoost/LightGBM); deep learning isn't necessarily stronger — Kaggle competition practices since the 2010s repeatedly confirm this.

Counterexample. "Blind deep learning" is the most common: small tabular data, hundreds of thousands of samples, yet deploying CNN/LSTM because "deep learning is state-of-the-art." The usual outcome: a model that's expensive, hard to explain, and not necessarily better than tree models. The opposite-direction counterexample: using simplicity as an excuse to refuse progress — evidence already shows tree models underfit, the business has absorbed larger capacity, yet sticking to "we only use linear models." Occam's razor is "don't add entities unnecessarily," but when necessity arrives, it's time to add.

Complexity additions must have cost awareness

Every new complexity term (new model family, larger embedding dimensions, longer training) goes in the "experiment ledger": how much metric gain did it bring? How much training/inference/ops/explainability cost did it cost? If gains can't be quantified, don't add by default.

Principle 3: Evaluation Design Before Modeling ​

Why. How good a model is is entirely defined by the evaluation protocol: how data is split, what metrics are used, who it's compared against, how "significant difference" is judged. Without a protocol, experiment conclusions are castles in the air — you tune for ages, only to find the split method is entirely unsuited to the business's time structure; you only look at accuracy, completely ignoring that this metric is meaningless under class imbalance. Evaluation-first also has a psychological effect: it forces you to think clearly about "what counts as success" before touching code — this matters more for project fate than any model choice. See Model Evaluation and Validation and What Evaluation Really Means in Practice.

How to do it. Before writing model code, write an evaluation protocol (even if just a note):

  1. Splitting method: random split? time-based split? by-group (user/article) split? Time-series problems almost always need time-based split — random splitting lets future information seep into the training set, see "temporal leakage" in Common Pitfalls.
  2. Metrics: one primary metric, two or three auxiliary. Metrics should map to business (Principle 10), and specify how to compute under imbalance and multi-class.
  3. Baseline: use the baseline numbers defined by Principle 1.
  4. Significance and variance: on small datasets, a single split's metric is meaningless; use K-fold CV or repeated experiments to report mean and standard deviation.
python
from sklearn.model_selection import TimeSeriesSplit  # use time-based split for time-series scenarios
from sklearn.metrics import f1_score

# Evaluation protocol written in advance:
# - Split: TimeSeriesSplit(n_splits=5)
# - Metric: F1 (macro average), with ROC-AUC and per-class recall as auxiliaries
# - Baseline: majority-class F1 = 0.23 (recorded)
# - Judgment: new model must exceed baseline by 0.05 on all 5 folds, and worst fold must not fall below baseline

Counterexample. Classic triple: ① split learned after the fact — using train_test_split random split for time-series, evaluation numbers look beautiful but are meaningless; ② metrics self-selected — severe class imbalance but only looking at accuracy, getting a "always predict majority" fake champion; ③ single experiment decides the winner — ran one CV and announced "new model is better," not realizing this split might just be luck, flip the seed and the conclusion flips. The evaluation protocol should be written down and reviewed by colleagues before modeling, not "patched" after the model is trained.

Principle 4: Data Quality > Model Complexity ​

Why. ML has a harsh conservation law: the model is just a machine that extracts patterns from data, and the data's pattern boundary is the model's ceiling. A dataset with half-mislabeled samples teaches wrong patterns to any model; a feature with 60% missing rate can't be saved by any advanced architecture. Peter Norvig's observation in The Unreasonable Effectiveness of Data still holds: "more data usually beats smarter algorithms." The industry adage "garbage in, garbage out (GIGO)" isn't saying models don't matter; it's saying in most real projects, the ROI of the data stage far exceeds the model stage.

How to do it. Make "data health check" a fixed pre-modeling step, not something you remember to do:

Check ItemFocusTools/Methods
Missing ratePer-column missing proportion, determines imputation/deletion strategydf.isnull().mean()
Duplicate rateDedup by row, by entity — distinguish "true duplicates" from "multiple records per user"df.duplicated() + business judgment
Label qualityMislabel rate, label distribution, class imbalanceSample for manual review
Value range and cardinalityPer-column value distribution, outliers, category cardinalitydf.describe(), histograms
Time spanData's time coverage matches the prediction targetmin/max timestamps

Document health check results in experiment logs (Principles 5, 9), and use them as a reference for error analysis (Principle 8) — when something anomalous happens, first ask "is this a data problem or a model problem?"

Counterexample. Counterexample 1: severely mislabeled labels (an entire channel's labels were wrong), but the team spent three weeks tuning and adjusting architecture because "tuning models feels more technical." Counterexample 2: treating "cleaning" as "deleting rows" — directly deleting outliers, deleting the business's most valuable high-value customers. Counterexample 3: over-cleaning — deleting "looks weird" but genuinely existing patterns during the cleaning stage, "washing" the data into another distribution. See Data Engineering for the proper approach.

Principle 5: Reproducibility Is the Bottom Line ​

Why. An unreproducible experiment equals not having done it: three days later, you can't answer "how was that 0.89 run?"; teammates can't build on your results. Reproducibility isn't just a scientific integrity issue; it's the prerequisite for collaboration and iteration — without it, everyone gets unique results in their "unique environment," and discussion and merging are impossible. The three big culprits of unreproducibility in ML: randomness (initialization, sampling, shuffle, multi-GPU parallelism), dependency drift (library version upgrades), data drift (data files overwritten or regenerated).

How to do it. A "minimum reproducible set" has four pieces:

python
# ① Fix random seeds (after importing libs, before any random operations)
import random
import numpy as np

SEED = 42
random.seed(SEED)
np.random.seed(SEED)
# Also set in deep learning:
# torch.manual_seed(SEED); tf.random.set_seed(SEED)
  1. Random seeds: globally fixed and recorded in experiment config; note some algorithms (multithreading, GPU atomic ops) can't be fully deterministic — accept "reproducible to a certain number of decimal places."
  2. Dependency locking: requirements.txt down to version numbers (or lock file / Docker image); pip freeze results archived with the experiment.
  3. Data versioning: data files hashed (sha256) and recorded, or incorporated into DVC or similar data version management; record the data version used per experiment.
  4. Experiment records: model code + hyperparameters + data version + metrics + environment version — five elements present for a valid experiment.

Don't overestimate "fixed the seed" issue

Fixing the seed is the minimum requirement. True reproducibility also includes: data split code and modeling code in the same repo, the same code producing the same data, evaluation scripts from the same source with training scripts. In other words, code is documentation, documentation is code.

Counterexample. Counterexample 1: an experiment that ran 0.91 two weeks ago, today's reproduction gives 0.84, because numpy was upgraded from 1.21 to 1.26 during the interim, and nobody recorded it. Counterexample 2: hyperparameters scattered in Jupyter's third cell, changing cell order changes the result set. Counterexample 3: downloaded a new data file from a cloud drive (content updated), old experiments all "drift." Engineering solutions for these issues are fully discussed in MLOps.

Principle 6: Always Prevent Data Leakage ​

Why. Data leakage is the number one cause of "inflated offline metrics, immediate crash on deployment," and the most insidious error in ML — because it looks perfectly normal: training converges, validation metrics are beautiful, the model seems "very smart." The truth is the model memorized answers into its parameters: it used information at training time that won't be available at inference. Every percentage point you celebrate at offline metrics will be returned with interest on deployment. Google's Rules of ML warns: "train-serve bias and data leakage are the two most common types of system design errors."

How to do it. Three disciplines:

  1. Split first, process later: dedup, normalize, impute, encode — all must fit on the training set first, then transform other sets; use sklearn.pipeline.Pipeline to guarantee by design.
  2. Split respects time and entities: time-series problems use TimeSeriesSplit or walk-forward; records of the same entity (same user, same article) must never be distributed across sets.
  3. Three questions per feature: ① Can this feature be obtained at prediction time? ② Is it produced after prediction time? ③ Is it determined by the same event as the target? Any answer of "no / yes / yes" means removing from the feature table.
python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,
                                                    random_state=42)
# ✅ Split first; scaler only fits training set
pipe = Pipeline([("scale", StandardScaler()),
                 ("clf", LogisticRegression())])
pipe.fit(X_train, y_train)

Counterexample. Counterexample 1: global dedup before splitting, rewritten versions of the same article fall on train and test sides; the model "memorized" the answers. Counterexample 2: predicting "whether a customer churns" using "number of complaints this month" as a feature — but complaints mostly occur after churn; this is target leakage. Counterexample 3: computing mean on full data to impute missing values, test information smuggled in through the fill statistics. These four words (dedup, impute, scale, time) — every detail deserves a separate article — see "data-level pitfalls" in Common Pitfalls and Anti-Patterns.

Principle 7: Deployment Must Include Monitoring ​

Why. Model deployment isn't the end; it's the moment it truly starts to rot. After deployment, three things happen: ① data drift — the user distribution changes (new user groups, new seasons, policy changes), training distribution and online distribution drift apart; ② concept drift — the patterns themselves change ("click = interest" no longer holds after a certain point); ③ system drift — the upstream data source of the feature pipeline changed format, what's fed in is no longer what the model recognizes. Without any monitoring, these three things can quietly happen for months until a business stakeholder asks "why is the recommendation so bad" and you discover.

How to do it. The deployment checklist must include at least four monitoring dimensions:

Monitoring DimensionWhat to WatchWarning Signal
Data driftDistribution of key features (mean, quantiles, category proportions) vs. training distributionDistribution shift beyond threshold / PSI spikes
Prediction distributionDistribution of model predictionsPrediction mean drifts, extreme predictions increase
Business metricsConversion rate, CTR, retention, etc.Significant month-over-month drop in metrics
System healthLatency, QPS, error rate, feature missing rateLatency rises, missing rate spikes

Monitoring is proactive: not just "send an alert when it breaks," but set reasonable thresholds and auto-inspections. Deploy gradually via canary/shadow mode (shadow → canary → full), checking monitoring metrics at each step before proceeding. See MLOps.

The ultimate purpose of monitoring

Monitoring isn't for you to "watch the model die"; it's to feed the iteration loop (Principle 8): distribution drift and error patterns found online are the best input for the next round of data cleaning, feature engineering, and retraining. See the full model lifecycle process in MLOps.

Counterexample. Counterexample 1: no monitoring dashboard after model deployment; recommendation quality quietly collapses three months later, discovered by the business side before engineers. Counterexample 2: only monitoring server latency (system health), completely ignoring prediction distribution and business metrics — the server is alive, but the model no longer "recognizes" the world. Counterexample 3: monitoring exists but thresholds are guessed, daily false alarms, team develops learned helplessness to alerts.

Principle 8: Iteration Driven by Error Analysis, Not Feelings ​

Why. The trap in ML iteration is: there are infinitely many directions for improvement, but resources are finite. Switch model architecture? Add features? Add data? Clean data? Adjust threshold? Each might work, but only evidence tells you which path to take. Error analysis is that evidence chain: collect the samples the model predicts wrong, classify them by cause, and let the errors themselves tell you whether the bottleneck is data, features, or model capacity. Andrew Ng repeatedly emphasizes in Machine Learning Yearning that: teams that intuitively allocate optimization work typically waste a lot of time on problems that error analysis would have solved immediately.

How to do it. A standard loop:

  1. Collect errors: take predicted-wrong samples from the validation set (or real failure samples sampled from online monitoring).
  2. Bucket: classify errors by cause — labels themselves wrong, insufficient information to judge, rare categories, missing features, ambiguous boundaries... one bucket per category.
  3. Quantify: count the proportion of errors in each category, sort by proportion.
  4. Targeted fix: prioritize fixing the largest error bucket — data problem → clean / add samples for that category; feature problem → add features; insufficient expressive power → consider switching models.
  5. Regression validation: after fixing, re-run the evaluation protocol (Principle 3), confirm no new error buckets were introduced.
1000 error samples
├── Label mislabeled     312 samples ──→ Data cleaning (highest priority)
├── Rare category missed 245 samples ──→ Add samples / category reweighting
├── Feature missing      178 samples ──→ Add feature pipeline
├── Boundary ambiguous   154 samples ──→ Accept, adjust decision threshold
└── Other                111 samples ──→ Defer

Counterexample. Counterexample 1: AUC goes from 0.80 to 0.82, team announces "new model wins," but nobody has looked at a single predicted-wrong sample — the 0.02 gain might just be overfitting the validation set. Counterexample 2: "deep learning should be better" by gut feeling, switching architecture and tuning for three weeks, performance actually gets worse — because the real bottleneck is label quality (error analysis tells you immediately). Counterexample 3: treating error analysis as a one-time action rather than a loop — error analysis isn't warmup before tuning; it's the starting point of every iteration.

Principle 9: Write Docs and Comments ​

Why. ML's biggest hidden cost is handoff cost: the person looking at your code three months from now (often yourself) needs to figure out: why this model? why were features computed this way? why is the threshold 0.5 and not 0.7? These "whys" aren't in the code; they're only in the mind at the time — without writing, they're gone forever. The most expensive thing in a team isn't training GPUs; it's re-understanding things already done.

How to do it. Three categories of documents, each with its role:

DocumentContentLifecycle
Project READMEProject goal, data source, evaluation protocol, how to runEntire project
Experiment logEvery experiment: goal, baseline, changes, metrics, conclusionsIterates with experiments
Decision record (ADR)Key choices: why this model / feature / threshold, what alternatives were consideredPermanent

The value of code comments is in "why" not "what":

python
# Why log1p compression on this feature here: the feature has a long-tail distribution,
# feeding directly to LR would make high-value samples dominate gradients (verified in 2024-03 experiment A3)
X["amount"] = np.log1p(X["amount"])

Counterexample. Counterexample 1: experiment folder has only final_v2.ipynb, 30 cells with no comments, the most critical cell (data split) has been manually edited three times with no record — nobody (not even the author) dares to guarantee this is the code that produced 0.89. Counterexample 2: docs and code decoupled — README describes old split logic, code has switched to TimeSeriesSplit; the doc becomes misleading instead. Docs must change alongside code — this is Principle 5 (reproducibility) extended to the collaboration level. Terminology alignment starts at Glossary.

Principle 10: Understand Business Goals ​

Why. More than half of ML project failures die from problem definition errors: offline metrics are optimized beautifully, but the business gets no benefit, or is even harmed. A classic case is recommendation systems optimizing "click-through rate": CTR goes up, but users find content doesn't match expectations, bounce rates skyrocket, trust drops — offline metrics won, business lost. The opposite classic case: risk control models optimizing "recall" while ignoring false alarm cost, blocking 90% of normal users, and the business loss far exceeds the gain from intercepted fraud. Metrics aren't the business itself; metrics are the quantifiable proxy for the business, and proxies distort.

How to do it. Before modeling, walk through the "business → ML" translation chain:

Business goal (e.g., improve repeat purchase rate)
   → Business decision (e.g., which users get a 10% discount coupon)
   → Needed prediction (e.g., predict user's 30-day repeat purchase probability)
   → Evaluation metric (e.g., ranking ability AUC of the prediction / yield calibration)
   → Loss function (e.g., binary cross-entropy, weighted)

Every link must be able to answer "why." In particular, ask three questions: ① After this prediction goes live, who makes what decision? ② What's the cost of making the wrong decision (asymmetric costs should be reflected in metrics and thresholds)? ③ When the metric goes up, does the business truly improve? These three questions are fully unpacked in Overall Architecture Anatomy.

More metrics isn't better

There's always only one primary metric; the rest are guardrails. Multiple primary metrics mean blurred goals, and the team will pick the most flattering one to "win." The right approach: one primary metric + two or three guardrail metrics that must not degrade (e.g., CTR goes up, bounce rate doesn't worsen).

Counterexample. Counterexample 1: the business says "we want an intelligent customer service," and the engineer directly starts training a dialogue model, nobody confirmed whether the real pain point is "answer accuracy" or "response latency" or "headcount cost." Counterexample 2: only looking at offline metrics (AUC, F1), no business success standard defined post-deployment, project "success" or "failure" entirely by gut feeling. Counterexample 3: metrics and business proxy distort without self-awareness — optimizing "video completion rate" causes the recommendation system to only push 30-second shorts, destroying the long-video ecosystem. ML is finding patterns from data, but whether patterns serve the business is a design-level matter.

III. Exceptions to the Principles: When You Can Break Them ​

Principles aren't dogma. Two scenarios where you can intentionally, controlledly relax some principles:

1. Rapid PoC (Proof of Concept). The goal is to answer "does this path work," typically with a 1–2 week deadline, producing a conclusion rather than a production system. At this point, you can relax: full monitoring (Principle 7), complete docs (Principle 9), even partial reproducibility (Principle 5).

2. One-time research / offline analysis. No deployment, no iteration, purely answering analytical questions (e.g., "is this feature useful"), relaxing monitoring and business loops is reasonable.

But exceptions have boundaries — here are "non-negotiable" red lines:

ScenarioCan RelaxCannot Relax
Rapid PoCMonitoring, docs, reproducibilityBaseline (Principle 1), anti-leakage (Principle 6)
One-time analysisMonitoring, iteration loopEvaluation protocol (Principle 3), anti-leakage (Principle 6)
Formal productionNothing (all apply)All apply

Why can't anti-leakage and baselines ever be skipped? Because a PoC's conclusions are only credible if its evaluation is credible. A leaked PoC produces the wrong conclusion "this path works," leading the entire team into a ditch; a PoC without baselines can't answer "how much better this solution is than the status quo." Engineering disciplines like monitoring and docs can indeed be simplified in "run-and-dump" scenarios.

Put a "disclaimer" on exceptions

When breaking principles, write in white and black: "This project is a PoC; conclusions cannot be directly used for production decisions; to go live, must complete X, Y, Z." This allows speed while preventing PoC conclusions from quietly becoming production conclusions.

Counterexample (exceptions abused). Most common: "PoC used as production" — PoC works, then opens to all users directly without any production workflow: no monitoring, no canary, no rollback plan. Second: "research exemption" indefinitely extended — the project starts to be deployed but still uses "this is research" to exempt all discipline. The exception's validity period should be defined by the goal: once feasibility is confirmed, PoC status ends immediately.

IV. Team-Level Practices ​

Principles work when a single person executes them, but to persist in a team, they need supporting mechanisms. Four experiences:

1. Model Code Review ​

Code review is standard in traditional software engineering, but often absent in ML projects — because "it runs" masks too many problems. Give model reviews a specialized checklist:

  • [ ] Was data split before feature engineering (anti-leakage)?
  • [ ] Are scaling / imputation / encoding only fit on the training set? Are they chained in a Pipeline?
  • [ ] Are random seeds fixed? Are dependency versions locked?
  • [ ] Do metric definitions match the evaluation protocol? Do they correspond to business goals?
  • [ ] Are experiment records complete (data version, code version, hyperparameters, metrics)?
  • [ ] Are there features "available at training but not at inference"?

2. Experiment Standards and Shared Baselines ​

The team maintains a shared baseline repo: a baseline code + a baseline data set + a set of authoritative baseline metrics. Any new experiment builds on this, preventing "each person has their own baseline." Experiment records use a unified template, at minimum including:

Experiment ID: EXP-014
Goal:     Verify the lift on repeat purchase prediction by adding the "user last 30-day activity" feature
Baseline: EXP-009 (logistic regression, AUC 0.812)
Changes:  Feature +1 (act_d30); model/hyperparams unchanged
Result:   AUC 0.818 (5-fold, ±0.004); click feature importance Top 3
Conclusion: Accept, include in next round

3. Unified Metric Definitions ​

"I say F1 you mean macro average, he says accuracy counts top-1" — in cross-team collaboration, metric definition disagreements make all comparisons meaningless. Extract metric calculation into a team-shared function library (one codebase, all teams reference), and define each metric's applicable scenarios in docs. This is Principles 5 and 9 at team level.

4. Feedback Loop from Monitoring to Iteration ​

The team-level process's end form is a closed loop: online monitoring detects drift → error analysis locates cause → data / feature / model improvement → new experiment → canary deployment → back to monitoring. Write into team policies: "who is responsible for monitoring post-deployment, how often to review, what signals trigger retraining" — instead of relying on one person's sense of responsibility.

        ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
        │ Online monitor │───▶│ Error analysis│───▶│ Data/feature fix │
        └──────────────┘    └──────────────┘    └──────────────┘
              ▲                                        │
              │                                        ▼
        ┌──────┴───────┐    ┌──────────────┐    ┌──────────────┐
        │ Canary/shadow │◀───│ Experiment +  │◀───│ Re-training  │
        │ deployment    │    │ review        │    └──────────────┘
        └──────────────┘    └──────────────┘

The psychology of team discipline

Design principles fail in teams rarely because the principles are wrong, but because they haven't become default processes. Put the evaluation protocol into PR templates, make experiment records required-issue fields, make monitoring dashboards a deployment prerequisite — once principles go from "self-awareness" to "mechanism," they no longer depend on everyone's willpower. The engineering foundation of reproducibility and version management is in MLOps.

V. Further Reading ​

Continue on-site:

External references (all real public resources):

  • Google. Rules of ML — the most classic ML engineering rule set, the source of many principles in this article
  • Google. Machine Learning Crash Course — official intro course covering engineering essentials like data leakage, train/serve bias
  • scikit-learn official docs: Model selection — splitting, cross-validation, grid search authoritative reference
  • scikit-learn official docs: Pipeline — the standard tool to prevent statistic leakage by design
  • Kaggle Learn. Data Leakage — interactive anti-leakage practice
  • Andrew Ng. Machine Learning Yearning (publicly free) — classic exposition on error analysis and priority-setting
  • Provost & Fawcett. Data Science for Business (O'Reilly) — classic textbook on the relationship between metrics and business goals
  • Peter Norvig. The Unreasonable Effectiveness of Data (2009) — classic exposition on "data > models"