Theme
What Is Machine Learning
One-sentence definition: Machine Learning is the technology that enables computers to automatically discover patterns from data and use those patterns for prediction or decision-making—programmers don't write "rules" directly, but instead write "programs that learn rules," letting models find answers in the data themselves.
The most "machine learning" code you've ever written might be a single line: LinearRegression().fit(X, y). Unpack that line: fit doesn't make the program execute a pre-written algorithmic flow—instead, it makes the program solve for a set of parameters from data X and labels y itself. The starting point of the entire machine learning discipline is the paradigm shift behind this line of code: from "humans writing rules" to "data generating rules."
I. Definition: Understanding from Three Perspectives
1. Practical Level: What Does It Do
A machine learning system takes inputs (features) and produces outputs (predictions/decisions), but its uniqueness lies in how the rules for producing outputs are derived:
Traditional programming: Rules (written by programmer) + Data ──→ Answer
Machine learning: Data + Answer (labels) ──→ Rules (learned automatically by model)Traditional programming is "program = algorithm + data structure," where the programmer enumerates all possible cases to write judgment logic. Machine learning is "program = learning algorithm + data," where the programmer only designs the learning algorithm and evaluation criteria, and the specific rules (model parameters) are determined by data. Any problem where "rules are hard to write by hand, but samples are easy to obtain"—image recognition, speech recognition, natural language, recommendation ranking—is machine learning's domain.
2. Academic Level: Tom Mitchell's Classic Definition
The most cited definition in the machine learning field comes from Tom Mitchell (1997):
"A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P, if its performance at tasks in T, as measured by P, improves with experience E."
The power of this definition is that it provides four quantifiable elements of machine learning:
| Element | Meaning | Example (spam classification) |
|---|---|---|
| Task T | What needs to be done | Determining whether an email is spam |
| Experience E | Data used for learning | 100,000 labeled emails |
| Performance P | Metric for measuring quality | Classification accuracy, F1 |
| Improvement mechanism | Learning algorithm | Updating model parameters to increase P |
Any problem that can be characterized by these four elements—task, experience, performance measure, and improvement mechanism—is suitable for machine learning; without any one element, the problem is either rule-based programming or an open-ended problem that cannot be learned. This framework is also the source of all subsequent evaluation methods—see Model Evaluation and Validation.
3. Mathematical Level: The Function Fitting Perspective
Mathematically, the vast majority of machine learning does the same thing:
Given an unknown target function f, use training samples (xᵢ, yᵢ) to approximate f(xᵢ) = yᵢ, learn a hypothesis function ĥ such that ĥ(x) ≈ f(x), and hope that ĥ gives reasonable outputs for unseen inputs.
"Performing well on unseen inputs" is called generalization. The core problem of all machine learning theory—the bias-variance tradeoff, regularization, cross-validation, overfitting—revolves around generalization. This explains why machine learning is not "memorizing data": a model that memorizes the training set perfectly but performs terribly on new data (overfitting) is professionally equivalent to not having learned at all. See Overfitting and Regularization for details.
Why can machines "learn"?
The fundamental prerequisite for machine learning to work is the existence of reusable statistical patterns in data. As long as the same pattern (e.g., "emails with 'winning' links are mostly spam") holds true for both samples and future data, the model can learn from samples and transfer to the future. This assumption is called the independent and identically distributed (i.i.d.) assumption—it doesn't always hold (data drift), but it is the foundation of virtually all machine learning methods.
II. Three Modeling Paradigms
By the "form of experience E," machine learning is divided into three paradigms:
| Paradigm | Form of Experience | Goal | Representative Methods |
|---|---|---|---|
| Supervised Learning | Labeled samples (x, y) | Learn the x→y mapping | Linear/logistic regression, decision trees, SVM, neural networks |
| Unsupervised Learning | Unlabeled samples x | Discover intrinsic data structure | Clustering, dimensionality reduction, association rules, autoencoders |
| Reinforcement Learning | Reward signals from environment interaction r | Learn a policy that maximizes cumulative reward | Q-learning, policy gradients, AlphaGo |
Their differences lie not in algorithmic complexity, but in "where the answer comes from": in supervised learning, someone provides the ground truth (labels); in unsupervised learning, there is no answer, only the structure inherent in the data itself; in reinforcement learning, there isn't even an answer—only a reward signal given after the fact ("this step was good/bad"). See Supervised Learning, Unsupervised Learning, and Reinforcement Learning for details.
You can remember the relationship among them this way: supervised learning learns "judgment," unsupervised learning learns "structure," and reinforcement learning learns "decision sequences." Deep learning is not a fourth paradigm, but a powerful family of tools for implementing the first three—using multi-layer neural networks for function fitting, it can be applied to supervised tasks (image classification), unsupervised tasks (autoencoders), and reinforcement tasks (DQN).
III. Machine Learning vs. Adjacent Concepts
This is the most confusing part for beginners, so let's clarify each one:
| Concept | Meaning | Relationship to ML |
|---|---|---|
| Artificial Intelligence (AI) | The overarching goal of making machines behave intelligently, encompassing reasoning, planning, perception, language, learning, etc. | Machine learning is a subfield of AI (and currently the most successful one); rule-based systems, knowledge graphs, and search also count as AI but not as machine learning |
| Deep Learning (DL) | A family of technologies using multi-layer neural networks for machine learning | A subset of machine learning; beyond deep learning, there are many classic methods (tree models, SVM, Bayesian) |
| Traditional Programming | Programmers hand-write rules | The opposite of machine learning; but in real systems, the two complement each other (rules as fallback + model as core) |
| Statistics | The discipline of inferring populations from data | A theoretical close relative of ML; statistics emphasizes inference and causality, while machine learning emphasizes prediction and scale |
| Data Science | The complete workflow of using data to support decisions (including visualization, dashboards, A/B testing) | Machine learning is the modeling step within the data science workflow; data scientists also do a lot of non-modeling work |
| Optimization | The mathematical branch for finding extrema of objective functions | The engine of machine learning; training models is essentially solving an optimization problem |
Two points worth expanding on:
Machine Learning vs. Deep Learning: This is not a "legacy vs. new" replacement relationship. Deep learning crushes classic methods on unstructured data like images, speech, and text, but on tabular data, tree models (XGBoost, LightGBM) remain often stronger—practice in Kaggle competitions throughout the 2020s repeatedly confirms this. Mature teams select models based on data type rather than "blindly going deep learning." See Tree Models and Ensemble Learning and Deep Learning Fundamentals.
Machine Learning vs. Statistics: The two share many tools (regression, hypothesis testing, Bayesian), but their temperaments differ: statistics emphasizes inference (is this effect significant? what's the confidence interval?), while machine learning emphasizes prediction (minimizing error on new data). Practically, machine learning relies more on compute and data, while statistics relies more on experimental design and modeling of data-generating mechanisms. Modern practitioners need both perspectives—relying only on the statistical view ignores scaling capabilities, while relying only on the ML view ignores the fundamental question of "how did your data come to be?"
IV. A Minimal Machine Learning System: Linear Regression Step by Step
Strip away all library abstractions, and a complete "machine learning workflow" needs only four steps. Using house price prediction as an example (feature: area x, label: price y):
python
import numpy as np
# ── ① Data: samples (x, y) ─────────────────────────────
X = np.array([50, 60, 70, 80, 90, 100]) # area (m²)
y = np.array([120, 150, 175, 210, 240, 270]) # price (10k CNY)
# ── ② Model: hypothesis function ĥ(x) = w·x + b ────────
# The model defines "what it can look like"; parameters (w, b) to be learned
# ── ③ Learning algorithm: gradient descent to minimize loss ──
# Loss function L(w,b) = 1/n Σ (yᵢ - (w·xᵢ + b))²
def gradient_descent(X, y, lr=0.001, epochs=1000):
w, b = 0.0, 0.0
n = len(X)
for _ in range(epochs):
pred = w * X + b
dw = (-2/n) * np.sum(X * (y - pred)) # gradient of loss w.r.t. w
db = (-2/n) * np.sum(y - pred) # gradient of loss w.r.t. b
w -= lr * dw # update along negative gradient
b -= lr * db
return w, b
w, b = gradient_descent(X, y)
print(f"Learned rule: price ≈ {w:.2f} × area + {b:.2f}")
# ── ④ Evaluation: predict on new data ──────────────────
new_x = 85
print(f"Predicted price for 85m²: {w * new_x + b:.1f} (10k CNY)")Just these four steps: data → model → learning algorithm → evaluation. These four steps form the skeleton of any machine learning project and are the starting point from which Anatomy of an ML System unfolds. Note that in this example, the programmer didn't write "the price rule for an 80 m² house"—the rule was learned by the model from the data. That's the whole secret of machine learning.
Try it yourself
Run this example in Jupyter, modify the loss function (try absolute error instead), change the learning rate, and observe the convergence curves of w and b. Ten minutes of hands-on experimentation beats reading the concepts ten times. For a more complete workflow, see Building an ML Project from Scratch.
V. Why Machine Learning Works: Three Empirical Milestones
"Machine learning works" is not a matter of faith—it's a fact that can be verified with data. Three milestone-level demonstrations:
1. The Perceptron Lights the First Flame (1957–1969)
In 1957, Frank Rosenblatt invented the Perceptron, demonstrating on a simulator that machines could "learn from samples," sparking the first AI wave. In 1969, Minsky and Papert proved in Perceptrons that a single-layer perceptron couldn't even represent XOR, causing the wave to cool quickly—a lesson that still holds today: the match between model capacity and problem complexity determines a technology's lifespan.
2. Deep Learning Surpasses Human Baseline (2012–2015)
In 2012, AlexNet at the ImageNet classification competition (ILSVRC) brought the top-5 error rate down from 26.2% to 15.3%—nearly 10 percentage points below second place—igniting the deep learning revolution. In 2015, a 152-layer ResNet pushed the top-5 error rate to 3.57%—for the first time surpassing human-level performance (about 5.1%). This was the first widely recognized data point for "machines surpassing humans on perceptual tasks." See CNNs and Computer Vision and Classic Papers Deep Dive for details on CNNs and ResNet.
3. Scaling Laws for Language Intelligence (2018–2023)
In 2018, BERT (340M parameters) swept 11 NLP benchmarks. In 2020, GPT-3 (175B parameters) demonstrated the remarkable zero-shot/few-shot capabilities of large-scale language models. At the end of 2022, ChatGPT brought conversational large models to the public, and in 2023, GPT-4 reached the top 10% level on multiple professional exams. For the complete narrative of this history, see Evolutionary History and Large Language Models.
These three milestones reveal two paradigm shifts in machine learning: from "feature engineering" to "representation learning" (deep learning automatically learns features), and from "task-specific" to "universal foundation models" (pre-trained large models can be adapted to hundreds of tasks). Understanding these two shifts is understanding the direction of the entire field today.
An honest counterpoint
Machine learning's effectiveness is conditional: biased data yields biased models, scarce data means nothing to learn from, and distributional drift causes model degradation. In 2018, Amazon shut down a recruiting model that systematically discriminated against women (the model learned the bias of "male preference" from historical resumes)—a classic case of "biases in data amplified by machine learning." Machine learning is not a magic wand; it is the same tool that amplifies both signal and noise in data. See Interpretability and Fairness for discussions on interpretability and fairness.
VI. Tradeoffs and Tensions
Four tensions that beginners in machine learning encounter first:
- Model complexity vs. generalization: The more complex the model, the stronger its ability to fit training data, but the more likely it is to memorize noise and lose generalization. This is the core of the bias-variance tradeoff and the reason regularization exists.
- Interpretability vs. predictive power: Linear models are transparent but limited in expressive power; deep models are powerful but act like black boxes. Which you choose depends on the decision context: risk control and healthcare require explainability, while recommendation ranking prioritizes performance. See Interpretability and Fairness.
- Simple and effective vs. cutting-edge: Kaggle practice repeatedly proves that starting with simple baselines (linear regression, logistic regression, single trees) and gradually upgrading is always the optimal path. Jumping straight to the most complex model is the most common waste for beginners. See Design Principles.
- Technology vs. business: No matter how high the offline metrics (accuracy, AUC), if there's no improvement in business metrics (conversion rate, cost savings) after deployment, the project has failed. Half of a machine learning project's work is in problem definition and evaluation design; the other half is modeling.
VII. Where to Start: This Site's Learning Path
Machine learning sits at the intersection of engineering, mathematics, and business. This site is organized in the order of "build concepts → understand mechanisms → dissect cases → hands-on practice":
- Reading guide (you are here): Continue reading ML vs AI vs Deep Learning vs Data Science to clarify conceptual boundaries, use Evolutionary History to build a timeline, and use Anatomy of an ML System to build a full-site map.
- Core knowledge: Build your foundation with Supervised Learning → Model Evaluation and Validation → Overfitting and Regularization → Feature Engineering → Optimization and Gradient Descent, then add advanced modules like Unsupervised Learning, Deep Learning Fundamentals, and MLOps.
- Case dissections: Ground abstract concepts in specific models—Linear Models, Tree Models, CNN, Transformer, Large Language Models.
- Hands-on practice: Building an ML Project from Scratch, Design Principles, Common Pitfalls.
- Reference anytime: Glossary, Math Primer, Datasets and Tool Archives.
References
- Tom Mitchell. Machine Learning (1997) — source of the classic "learning from experience E" definition
- Krizhevsky, Sutskever, Hinton. ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012) — AlexNet, the ignition point of the deep learning revolution
- He, Zhang, Ren, Sun. Deep Residual Learning for Image Recognition (CVPR 2016) — ResNet, top-5 error rate of 3.57%, first to surpass human baseline
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (NAACL 2019)
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020) — GPT-3 and scaling laws
- Rosenblatt. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (1958) — original perceptron paper
- Minsky & Papert. Perceptrons (1969) — classic proof of single-layer perceptron limitations
- Amazon scraps secret AI recruiting tool that showed bias against women (Reuters, 2018) — a landmark case of data bias amplification