Theme
Anatomy of an ML System
What exactly does a "machine learning project" consist of? Most beginners imagine it as "writing models and running training," but a real production-grade system is far more than that. This article breaks a complete ML system into nine stages, with one sentence per stage explaining "what it does, why it matters, and where the pitfalls are," and links to corresponding pages across the site. This page is the full-site map—after reading it, you'll have a sense of the skeleton of the entire field.
I. Big picture
text
┌───────────────────────── Machine Learning System ──────────────────────────┐
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ ① Problem │──▶│ ② Data │──▶│ ③ Feature│──▶│ ④ Model Selection │ │
│ │ Definition│ │ Pipeline │ │ Engineering│ │ Algorithms & │ │
│ │ Business→ML│ │ Collect │ │ Encode │ │ assumptions │ │
│ └──────────┘ │/Clean │ │ /Scale │ └──────┬───────────┘ │
│ └──────────┘ └──────────┘ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │
│ │ ⑨ Monitor │◀──│ ⑧ Deploy │◀──│ ⑦ Evaluate│◀──│ ⑤ Training & │ │
│ │ & Re-train│ │ Service │ │ & Iterate│ │ Tuning │ │
│ └──────────┘ └──────────┘ └──────────┘ │ Loss/Optimizers │ │
│ └──────────────────┘ │
│ │
│ ⑥ Engineering foundation: experiment management / feature store / model│
│ registry / CI-CD (running through everything) │
└───────────────────────────────────────────────────────────────────────────┘The nine stages aren't purely waterfall (in practice, you loop back repeatedly), but each stage has its own specialized knowledge and common pitfalls. We dissect them one by one below.
II. Dissecting the nine stages
① Problem Definition: business problem → machine learning problem
The starting point of every ML project, and also the stage most easily skipped. The core is to answer four questions:
- Is this an ML problem? If rules can be written, write rules—costs half as much. The criterion: rules are hard to write by hand, but samples are easy to obtain.
- What are the inputs and outputs? What are the features X, what is the target y, and which of classification/regression/ranking/clustering does this fall into.
- What does success look like? Offline metrics (accuracy, AUC) must align with business metrics (conversion rate, cost savings).
- Is the data available? Without data, everything else is built on sand.
Get the problem definition wrong, and the more you work on it later, the more you'll be wrong. A classic cautionary tale: spent three months building a "predicting customer churn" model to AUC 0.95, only to discover the business didn't need "predicting who will churn" but "intervening on who to keep."
② Data Pipeline: from raw data to usable data
- Collection: logs, databases, third-party APIs, public datasets;
- Cleaning: missing values, outliers, duplicates, format unification (see Data and Data Engineering for common cleaning pitfalls);
- Labeling: supervised learning needs labels, and label quality directly determines the model's ceiling—"garbage in, garbage out" (GIGO) is the hardest rule of this stage;
- Splitting: training/validation/test sets—the split must simulate the real prediction scenario (time series cannot be randomly split).
Key lesson: data leakage—test set information bleeding into the training set, causing inflated offline metrics and production failure. This is the #1 cause of ML engineering disasters. See Model Evaluation and Validation.
③ Feature Engineering: the model's "nutritional formula"
Feature engineering is "transforming raw data into numerical features the model can digest": missing value handling, categorical encoding (one-hot/target encoding), numerical scaling (standardization/normalization), feature combinations, and domain features. In the classic ML era, "feature engineering set the ceiling"; in the deep learning era, it's still relied on in half the use cases.
Watch out for two extremes: too few features lead to underfitting, too many/dirty features introduce noise and even overfitting. The principle is "start with a simple baseline, then add features based on evidence"—see Feature Engineering.
④ Model Selection: algorithms and assumptions
Choosing a model isn't "the newer the better," but matching the data type and problem scale:
| Data type | Recommended approach | Why |
|---|---|---|
| Tabular data, small sample size | Logistic regression / tree models (XGBoost/LightGBM) | Stable, fast, interpretable, fewer hyperparameters |
| Images | CNNs (ResNet family) | Local inductive bias |
| Text/sequences | Transformer | Long-range dependency modeling |
| Large-scale unlabeled data | Pre-training + fine-tuning | Transfer learning |
The full logic of model selection is in Supervised Learning and How to Choose Frameworks and Tools.
⑤ Training and Tuning: loss, optimizers, and hyperparameters
- Loss function: MSE/MAE for regression, cross-entropy for classification—the loss should align with business objectives (e.g., predicting prices, squared error is overly sensitive to outliers);
- Optimizers: tradeoffs among SGD, Adam, AdamW (learning rate, momentum, weight decay);
- Hyperparameters: learning rate, batch size, tree depth/count, network layers—hyperparameter search is covered in Hyperparameter Tuning;
- Regularization: L1/L2, Dropout, early stopping, data augmentation—the arsenal against overfitting, see Overfitting and Regularization.
Classic training failures: loss not decreasing, vanishing/exploding gradients, overfitting without realizing it—each has a corresponding troubleshooting checklist (see Common Pitfalls and Anti-Patterns).
⑥ Engineering Foundation: experiment management, feature store, model registry
Three essential pieces of infrastructure for production-grade teams, making the nine stages reproducible and roll-back-able:
- Experiment management (MLflow, Weights & Biases): record data version, code version, hyperparameters, and metrics for each experiment—otherwise you'll have no idea three months from now which experiment performed best;
- Feature store (Feast, Tecton): shared features for training and inference, preventing "training features differ from online features";
- Model registry (MLflow Model Registry): model versioning, approval, and canary releases.
Small projects without a foundation can start with folders + Git, but start keeping experiment records from day one—this is the first rule of Design Principles.
⑦ Evaluation and Iteration: let data speak, not gut feeling
- Metrics: accuracy/precision/recall/F1/AUC for classification, MSE/MAE/R² for regression, NDCG for ranking—metrics must align with business objectives; accuracy is a trap under class imbalance (a "garbage model" with 99% accuracy);
- Validation: K-fold cross-validation, hold-out validation, rolling validation for time series;
- Baseline: always run a simple baseline (majority class, linear model) first—models must beat the baseline to be meaningful;
- Error analysis: look at misclassified examples, group by error type—the right way to iterate is "add features/tune based on error analysis," not random hyperparameter tweaking.
See Model Evaluation and Validation and Building Model Evaluation from Scratch for the full methodology.
⑧ Deployment: serving models
- Offline vs. online: batch prediction (run once daily) vs. real-time API (latency-sensitive);
- Serving methods: online inference API (FastAPI/KServe), streaming prediction, embedded (mobile/edge);
- Performance: model compression (quantization, distillation), inference acceleration (ONNX, TensorRT);
- Scale: model monitoring, A/B testing, canary deployment.
The full engineering side of deployment (including the gap between offline evaluation and online performance) is in MLOps and Model Deployment.
⑨ Monitoring and Re-training: models "expire"
Deployment is just the beginning: data distributions drift (user behavior changes, seasons change), and model metrics will degrade over time. Monitor three things:
- Data drift (input distribution changes): detected via PSI, KS statistics;
- Concept drift (the x→y relationship changes): e.g., the "features → churn" pattern broke during the pandemic;
- Prediction quality (when delayed labels arrive): online re-evaluation of click rates, conversion rates.
When drift hits the threshold, trigger re-training (scheduled or on-demand). This stage is one of the main reasons "ML projects have high failure rates"—many models don't die in training, but from being abandoned after deployment.
III. Weight distribution and skills map across the nine stages
Different roles spend very different energy across these nine stages, directly mapping to the role landscape in Career and JD Analysis:
| Stage | Algorithm Engineer | ML Engineer | Data Scientist | Data Analyst |
|---|---|---|---|---|
| ① Problem Definition | ★★ | ★★★ | ★★★★★ | ★★★★ |
| ② Data Pipeline | ★★ | ★★★★ | ★★★★ | ★★★★ |
| ③ Feature Engineering | ★★★ | ★★★★ | ★★★★ | ★★ |
| ④ Model Selection | ★★★★★ | ★★★★ | ★★★ | ★ |
| ⑤ Training & Tuning | ★★★★★ | ★★★★ | ★★ | — |
| ⑥ Engineering Foundation | ★★ | ★★★★★ | ★★ | — |
| ⑦ Evaluation & Iteration | ★★★★ | ★★★★ | ★★★★ | ★★ |
| ⑧ Deployment | ★★ | ★★★★★ | ★ | — |
| ⑨ Monitoring & Re-training | ★★ | ★★★★★ | ★★ | — |
Advice for beginners: ① and ⑦ are must-master for everyone (problem definition + evaluation), ③ ④ ⑤ are algorithm fundamentals, and ⑥ ⑧ ⑨ are engineering-side bonuses. Pick your route based on Learning Paths.
IV. Four perspectives that run through the entire site
- Conceptual: 12 articles in Core Knowledge thoroughly explain the principles behind the nine stages;
- Case-based: 10 articles in Classic Cases show how real models land in practice;
- Paper-based: 6 articles in Paper Deep Dives take you back to the foundational literature;
- Practical: 8 articles in Practice Guides walk you through building all nine stages with your own hands.
How to use this page
- To quickly build a big-picture view: reading this article is enough;
- To dive into a specific stage: click the links to the corresponding pages;
- To get hands-on: go directly to Building an ML Project from Scratch and check which of the nine stages the demo covers and which are missing—what's missing is what you need to fill in.
Further reading
- What Is Machine Learning—the starting definition for this article
- Evolutionary History—why the nine stages look the way they do
- MLOps and Model Deployment—full expansion of stages ⑧ and ⑨
- Data and Data Engineering—the deep end of stage ②
- Model Evaluation and Validation—the methodology for stage ⑦
- Design Principles—10 rules for getting the nine stages right
- Common Pitfalls and Anti-Patterns—classic failure scenes for each of the nine stages