Skip to content

Classic Paper Deep Dives

Quick overview From the Perceptron to GPT-3, deep reads of nine classic papers that changed machine learning: background, mechanisms, experiments, implications, plus the three-pass paper-reading method and the common patterns behind nine papers, helping you build a paper-reading coordinate system.

Classic Paper Deep Dives ​

In one sentence: These nine papers strung together form an evolutionary history of "modern machine learning"—from the first spark of the 1958 Perceptron to the scaling law ignited by GPT-3 in 2020, each one solved a specific pain point that was blocking the entire field at the time, then changed the world.

1. Why These Nine Papers ​

Line up the nine papers chronologically and you'll see a complete evolutionary history:

1958  Perceptron        Neurons first "learn from samples"
2012  AlexNet           Deep learning revolution officially detonates
2014  GAN               Adversarial game of generative models debuts
2015  ResNet            Networks first "deep to 152 layers"
2016  AlphaGo           AI first defeats humans at the pinnacle of human intellect
2017  Transformer       Attention replaces recurrence and convolution
2019  BERT              Bidirectional pretraining + fine-tuning paradigm solidifies
2020  GPT-3             Scaling law arrives, large model arms race begins
2020  DDPM              Diffusion models reshuffle the generative AI table
PaperFirst AuthorVenueOne-Sentence Contribution
The PerceptronRosenblatt1958, Psychological ReviewThe first "learnable" artificial neuron model
ImageNet Classification with Deep CNNsKrizhevsky2012, NeurIPSGPU + deep convolution + ReLU/Dropout, igniting deep learning
Deep Residual LearningHe2016, CVPRResidual connections, making networks deep to hundreds of layers
Attention Is All You NeedVaswani2017, NeurIPSPure-attention architecture Transformer, replacing RNN
BERTDevlin2019, NAACLBidirectional pretraining + fine-tuning, sweeping NLP benchmarks
Generative Adversarial NetsGoodfellow2014, NeurIPSZero-sum game between generator and discriminator
Denoising Diffusion Probabilistic ModelsHo2020, NeurIPSStable add-noise-remove-noise generation, rewriting image generation
Mastering the game of GoSilver2016, NatureMCTS + deep networks, conquering Go
Language Models are Few-Shot LearnersBrown2020, NeurIPS175 billion parameters, scale is capability

Why these nine and not others? Three reasons:

  • Each represents a paradigm shift: from "humans write rules" to "data learns rules" (Perceptron), from handcrafted features to representation learning (AlexNet), from sequence modeling to attention (Transformer), from discriminative to generative (GAN/DDPM), from specialized intelligence to general foundation models (BERT/GPT-3). Understanding these nine shifts captures the most important ideological backbone of 60 years in this field.
  • They're all readable lengths: except AlphaGo, most are around 10 pages, methods and experiments focused on one core idea—far more readable than today's massive technical reports.
  • Citation counts are in the tens to hundreds of thousands: they are "anchors" that have been tested and cited by the world repeatedly. Once you understand them, any follow-up work has a reference system.

Before you read

I recommend first reading Start Here to understand the section's three routes, then pairing with Evolutionary History to build a timeline. Each deep dive follows a fixed four-step structure: problem background → core mechanism → experimental conclusions → implications for today, making it easy to take notes as you read.

2. The Perceptron: Everything Begins (Rosenblatt, 1958) ​

1. Problem Background ​

In the 1950s, the Turing test had just been proposed, and computers had only just learned arithmetic. The mainstream AI approach was symbolism: writing knowledge as logical rules and having machines reason by those rules. Rosenblatt proposed a completely opposite connectionist approach: the brain is made of billions of neurons, each doing nothing more than "weighting inputs, summing, and activating when exceeding a threshold"—if machines built a simplified model with this structure, could they learn judgment from samples on their own, without humans writing rules?

2. Core Mechanism ​

The Perceptron is history's first "learnable" artificial neuron:

Input x₁,x₂,...,xₙ ──weighted──▶ Σ wᵢ·xᵢ ──threshold activation──▶ Output y ∈ {0, 1}

The learning rule is exceptionally simple: predict wrong → change weights → direction follows the error.

Predict:  ŷ = 1 when Σ wᵢxᵢ > threshold, else 0
Update:   wᵢ ← wᵢ + η·(y − ŷ)·xᵢ      (η is learning rate)

If the samples are linearly separable, this rule is guaranteed to converge in a finite number of steps—this is the Perceptron Convergence Theorem (Novikoff, 1962 gave a rigorous proof). Note what "convergence theorem" means: this isn't a luck-based heuristic; it's a learning algorithm with mathematical guarantees. It's the source of all later "optimization guarantee" thinking.

3. Experimental Results ​

Rosenblatt demonstrated with a dedicated hardware device, the Mark I Perceptron (photoelectric scanner + analog circuits): the machine learned to distinguish triangles from circles in photos. This was sensational at the time—"machines learned on their own," not "humans wrote rules into machines." The paper was published in the top psychology journal Psychological Review because its motivation was neuroscience: the Perceptron is more of a simplified hypothesis for biological neuron learning than an AI model.

4. Implications for Today ​

  • Structural prototype: weighted sum + nonlinear activation—the skeleton runs through every neural network that follows, all the way to today's Transformer.
  • The value of theoretical guarantees: the Perceptron convergence theorem shows that simple rules + provable convergence can produce credible intelligent behavior. Modern deep learning rarely offers theoretical guarantees, but the Perceptron reminds us: what can be proved, endures longer.
  • The boundary of capacity: in 1969, Minsky & Papert proved in Perceptrons that a single-layer Perceptron can't even represent XOR, triggering the first AI winter. This tells us: no matter how popular a direction is, if model capacity doesn't match problem complexity, you hit a theoretical ceiling. Over twenty years later, multi-layer Perceptrons + backpropagation (Rumelhart, 1986) solved XOR, proving the problem wasn't with "neural networks" themselves but with depth.

The Perceptron's dual legacy

The Perceptron left two lessons that still hold today: first, borrowing structure from neuroscience and learning parameters from data is a route that has been repeatedly validated; second, honest testing of theoretical boundaries (Minsky & Papert's book, though it killed the hype, was rigorous mathematics) pushes the field forward more than blind optimism.

3. AlexNet: The Detonation Point of the Deep Learning Revolution (Krizhevsky, 2012) ​

1. Problem Background ​

Before 2012, image recognition relied on "handcrafted features + classifiers": researchers hand-designed features like HOG and SIFT, then fed them to SVMs. The best result on the ILSVRC-2012 competition was a top-5 error rate of about 26%. While deep convolutional networks were theoretically viable, they just couldn't "train"—not enough compute to support large networks; using sigmoid/tanh activation, gradients vanished once the network went deep; hundreds of thousands of parameters severely overfit on small datasets. AlexNet solved this problem at once with three pillars.

2. Core Mechanism: GPU, ReLU, Dropout—Three Pillars ​

         Pillar 1                 Pillar 2               Pillar 3
     GPU Parallel Computing    ReLU Activation         Dropout + Data Augmentation
     ┌──────────────┐          ┌──────────────┐          ┌─────────────────┐
     │  Two GTX580s  │          │ max(0,x)     │          │ Randomly drop half│
     │ Conv parallel │          │ No saturation│          │ of FC layer neurons│
     │ ~10-20x faster│  Gradient│ of positive  │          │ Random crop/flip/  │
     │ training      │  doesn't │ interval     │          │ color perturbation │
     │               │  vanish  │ 6x faster    │          │ Equivalent to ens. │
     └──────────────┘          │ than tanh    │          │                  │
                               └──────────────┘          └─────────────────┘
  • GPU: using two GTX 580s (3GB VRAM each), split the 8-layer network (5 conv + 3 FC, ~60 million parameters) across two GPUs for parallel training, turning "large networks" from infeasible to "trainable in days."
  • ReLU: ReLU(x) = max(0, x). Its gradient is constantly 1 in the positive region, unlike sigmoid which saturates, greatly alleviating the vanishing gradient problem in deep networks. The paper compared on CIFAR-10: ReLU reached the same error about 6x faster than tanh.
  • Dropout + data augmentation: 60 million parameters on 1.28M images inevitably overfit. Dropout randomly drops FC layer neurons at 0.5 probability (equivalent to training an ensemble of sub-networks); simultaneously using random cropping, horizontal flipping, and PCA color perturbation expanded the effective sample size by tens of times. The paper's conclusion is direct: "without Dropout, severe overfitting occurs."

3. Experimental Results ​

AlexNet won ILSVRC-2012 with a top-5 error rate of 15.3%, second place at 26.2%—a direct, cliff-edge crushing of nearly 11 percentage points, an unmistakable lead. Over the next three years, deep learning became the only mainstream in vision. By 2015, ResNet pushed the error rate down further to 3.57%. 2012 is therefore called "the year of deep learning."

4. Implications for Today ​

  • The three pillars are a universal prescription: compute, good activation functions, anti-overfitting measures—training any large model today still requires these three.
  • View engineering details dialectically: the paper's local response normalization (LRN) and overlapping pooling were later shown to have limited or no effect, but ReLU, Dropout, and data augmentation remain standard operations to this day. Distinguishing "what the author thought was important" from "what time verified as important" is a basic skill of critical reading.
  • Representation learning paradigm: features are auto-learned by networks, replacing handcrafted feature engineering—this is deep learning's fundamental victory over traditional methods.
  • For deeper mechanism details, see CNN & Computer Vision and Deep Learning Foundations.

A reminder about one number

15.3% is a "for that time, that place" number: change the data split, hardware, and evaluation script of that era, and the number vanishes. What's truly worth taking away is the mechanistic judgment of "why the three pillars are indispensable," not that SOTA.

4. ResNet: Making Networks Deeper (He, 2016) ​

1. Problem Background ​

AlexNet had 8 layers, VGG had 19, and people found that "deeper networks = higher accuracy." What happens if you keep going deeper? The answer was surprising: degradation. A 56-layer network had a higher training error on ImageNet than a 20-layer one—not just test error, but training error was higher too, meaning deep networks couldn't even "memorize the answers." The problem wasn't overfitting; it was optimization: vanishing/exploding gradients made deep networks "unable to learn."

2. Core Mechanism: Identity Shortcut ​

ResNet's solution is embarrassingly simple: instead of learning the target mapping H(x), let the network learn the residual F(x) = H(x) − x.

         x
         │
         ▼
   ┌─────────────┐
   │ Conv→BN→ReLU │
   │ Conv→BN      │ ──▶ F(x)
   └─────────────┘
         │                    │
         └───── identity +x ──┘
               │
               ▼
           F(x) + x ──▶ ReLU ──▶ Output

Why this "add one line" works so well:

  • Gradient highway: during backpropagation, gradients can "cut across" the identity shortcut and propagate directly, eliminating vanishing gradient concerns for deep networks.
  • Theoretical explanation of degradation: if the network wants to learn an identity mapping H(x)=x, the residual block only needs to learn F(x)=0—pushing weights to 0 is much easier than learning an exact identity transform. Therefore, deeper networks perform at least as well as shallower ones.
  • Synergy with BatchNorm: residual blocks pair with BatchNorm (2015, Ioffe & Szegedy) internally, taking training stability up another level.

3. Experimental Results ​

The 152-layer ResNet brought top-5 error down to 3.57% (single model 4.49%) on ILSVRC-2015—first exceeding human baseline (estimated at ~5.1%). On CIFAR, the authors even trained 1000+ layer networks and still achieved convergence. From AlexNet's 15.3% to 3.57%, error was reduced fourfold in two years—"machines surpassing humans at perception" became a reproducibly verifiable fact for the first time.

4. Implications for Today ​

  • Residual connections are one of the most important architectural ideas in deep learning: today's Transformers, U-Net, diffusion model backbones all use residual connections. They turned "deep" from a curse into a gift.
  • Diagnose before prescribing: the authors first confirmed via experiment that degradation was "not overfitting" (otherwise training error should be lower), then targeted the root cause. This methodology of using experiments to locate failure modes is even more worth learning than residual connections themselves.
  • Extensions of the identity mapping idea: skip connections are now widely used in U-Net (diffusion models) and various multi-scale architectures; "giving information a shortcut" is a universal design principle.
  • Connect to AlphaGo's implications: structural "shortcuts" and learning "shortcuts" (learning from human game records first) are two sides of the same idea.

Why degradation isn't overfitting

"Deeper = worse" most naturally attributes to overfitting. ResNet refuted this with one figure: the 56-layer training error is higher than the 20-layer's. Overfitting is "good on training, bad on test"; here it's "bad on training itself." This one-character difference determines completely different solutions—don't add regularization; add residual connections. Precisely defining the failure mode is the first step of innovation.

5. Transformer: Attention Replaces Recurrence (Vaswani, 2017) ​

1. Problem Background ​

In 2017, machine translation was ruled by RNNs (especially LSTMs), but recursive structures had two inherent shortcomings: first sequential computation—processing token t requires finishing t−1 first, no parallelism, extremely slow training on long sequences; second long-range dependency—the farther apart information is, the harder it is to transmit, with gradients vanishing or exploding. Bahdanau's 2014 attention mechanism alleviated the second problem, but recurrence remained the backbone. The Google team posed a radical question: what if we completely remove recurrence and convolution, relying solely on attention?

2. Core Mechanism: QKV Self-Attention ​

Transformer's entire magic condenses into one formula:

Attention(Q, K, V) = softmax( Q·Kᵀ / √d_k ) · V

  Q (Query)  : "What am I looking for"
  K (Key)    : "What am I"
  V (Value)  : "What I provide"

Each token generates its own Q, K, V; dot-product Q against all tokens' K to get "attention weights" (dividing by √d_k prevents dot products from becoming too large, causing softmax gradients to vanish); then weight V by the weights—each token can see any token in the entire sentence in one step, eliminating the long-range dependency problem.

Four key designs around it:

  • Multi-head attention: split attention into h "heads" for parallel computation, each head attending to different subspaces (some handle syntax, some handle reference, some handle position), then concatenate and project. Experiments in the paper show multi-head significantly outperforms single-head.
  • Positional encoding: without recurrence, there's no order information; positional information must be explicitly injected into inputs via sine/cosine functions.
  • Residual connections + LayerNorm: every sub-layer is wrapped in "residual + normalization," ensuring deep networks are trainable.
  • Masked attention: when the decoder predicts the next token, it masks future positions to prevent "peeking at the answer."
Encoder layer ×6 ──────────────────▶ Decoder layer ×6
┌────────────────┐            ┌────────────────┐
│ Multi-head self-attention  │  Masked multi-head self-attention
│   ↓ (residual+LN)          │  ↓ (residual+LN)
│ Feed-forward network       │  Encoder-decoder attention
│   ↓ (residual+LN)          │  ↓ (residual+LN)
└────────────────┘            │ Feed-forward network
                              │  ↓ (residual+LN)
                              └────────────────┘

3. Experimental Results ​

On WMT 2014 English-German translation, BLEU reached 28.4 (prior SOTA 26.8); English-French reached 41.8 (prior SOTA 39.2), while training cost was only a fraction of the best recurrent models—8 P100 GPUs for 3.5 days. Dual domination in both quality and efficiency let "pure attention" quickly replace RNN.

4. Implications for Today ​

  • Universal base: Transformer became the shared backbone for GPT, BERT, diffusion model backbones—the foundation of all 2020s generative AI. See Transformer Case Study.
  • The design philosophy of "fewer assumptions + more data": removing the sequence inductive bias of recurrence/convolution, replacing it with a more general mechanism for stronger scalability—this approach proved extremely effective, and explains why the "pretrained large model" route won.
  • Complexity legacy: self-attention is O(n²); long sequences (entire books, full videos) remain a research frontier—sparse attention, linear attention, FlashAttention, and other engineering optimizations continue to evolve.
  • Note its relationship with residuals: every Transformer layer wraps a residual connection—ResNet's answer is what made Transformer deep enough. Classic papers are like this: they build ladders for each other.

6. BERT: Bidirectional Pretraining Opens the Large Model Era (Devlin, 2019) ​

1. Problem Background ​

By 2018, "pretraining + fine-tuning" had proven effective: pretrained first on a massive unlabeled corpus, then fine-tuned on task data. But existing pretrained language models had structural defects: GPT was unidirectional left-to-right (could only see left context); ELMo was "shallowly" concatenating two-direction LSTM representations. And a large number of understanding tasks—cloze, coreference resolution, inter-sentence relation judgment—naturally require seeing both directions simultaneously. How to build a "deeply bidirectional" language model?

2. Core Mechanism: MLM + NSP ​

BERT (Bidirectional Encoder Representations from Transformers) untied this knot with two pretraining tasks:

Task 1: Masked Language Model MLM (solves "deep bidirectionality")
  Input:  I [MASK] machine learning. [MASK] is a discipline that auto-learns patterns from data.
  Goal:   Predict the two [MASK]ed words ("like"/"it")
  → The model is forced to use both left and right context simultaneously, achieving deep bidirection

Task 2: Next Sentence Prediction NSP (solves "inter-sentence relations")
  Input:  [CLS] Machine learning is interesting [SEP] It learns patterns from data [SEP]
  Goal:   Determine whether the second sentence is the true next sentence of the first
  → Injects capability for QA, reasoning, natural language inference, and other sentence-pair tasks

Architecture: multi-layer bidirectional Transformer encoder. BERT-base: 12 layers, 110M parameters; BERT-large: 24 layers, 340M parameters. Pretraining corpus: BookCorpus + English Wikipedia (~3.3 billion tokens). For fine-tuning, only add a task-specific output layer on top and use very few training steps to adapt to any downstream task.

3. Experimental Results ​

BERT-large swept SOTA on 11 NLP benchmarks:

BenchmarkBERT-largePrior Best
GLUE score80.572.8
SQuAD v1.1 (QA F1)93.290.9 (human 91.2)
SQuAD v2.0 (F1)83.180.1
SWAG (commonsense reasoning)86.383.0

Notably on SQuAD v1.1, F1 of 93.2 first surpassed human level 91.2, drawing widespread public attention to "AI reading comprehension surpassing humans."

4. Implications for Today ​

  • Pretraining + fine-tuning paradigm: massive unlabeled data learns universal representations, a small amount of labeled data adapts to tasks—this paradigm defined the late 2010s through the 2020s. All large models are its descendants.
  • The bifurcation and unification of bidirectional vs. unidirectional: BERT proved bidirectional understanding is stronger, GPT proved unidirectional generation is more natural. Today's decoder-only large models reunite both via "causal masking + attention." Understanding this "diverge then reunify" trajectory is a key thread for understanding large model history. See Large Language Models.
  • Creativity in task design: turning "cloze fill" into a pretraining task (MLM), turning "next sentence judgment" into a pretraining task (NSP)—the pretraining objective itself is a research object worth innovating, as T5 and RoBERTa later did.
  • Toward scaling laws: BERT with 340M parameters + 3.3B tokens, performance scaling synchronously with both—this directly pointed to GPT-3's two years later: "parameters, data, compute grow by power law."

A common BERT interview question

"Why can BERT be bidirectional but GPT can't?" The answer lands on MLM: generative language models must predict the next token in sequence, inherently unidirectional; BERT gives up generation, only does understanding, using masked prediction to "see the full text but block one word," thereby gaining bidirectional context. Architecture choice is fundamentally a training objective choice.

7. GAN: The Game Theory of Generative Adversarial (Goodfellow, 2014) ​

1. Problem Background ​

Around 2014, deep learning was blooming on discriminative tasks (image classification, speech recognition), but generation—creating realistic images from random noise—was still poor: VAE and other explicit generative models either produced blurry samples or required complex variational derivations. Goodfellow, then a Ph.D. student, proposed a revolutionary idea: instead of painstakingly designing a generative model, let two networks adversarially compete, co-evolving in a game.

2. Core Mechanism: min-max Game ​

GAN consists of two networks in a zero-sum game:

  • Generator G: eats random noise z, outputs fake sample G(z), whose only goal is to fool the discriminator.
  • Discriminator D: judges whether input is real data or G's forgery, outputting the probability "this is real."
        Random noise z
           │
           ▼
   ┌─── Generator G ───┐   Fake sample G(z) ──┐
   └───────────────────┘                      ▼
                                     ┌─── Discriminator D ───┐
   Real data x ─────────────────────▶│ Real? / Fake?          │
                                     └───────────────────────┘
    D wants to win: spot all fakes (real → real, fake → fake)
    G wants to win: make D judge G(z) as real
    Game equilibrium: D always outputs 0.5 (can't distinguish), G has learned the data distribution

The objective function is min-max:

min_G max_D  V(D,G) = E_x[log D(x)] + E_z[log(1 − D(G(z)))]

Training alternates: fix G, update D; fix D, update G; repeat. Theoretical analysis shows that under ideal conditions, the game converges to a point where G's probability distribution equals the real data distribution p_data. The most crucial insight: you don't need to explicitly write the distribution density—just having a "scoring signal," the generator can implicitly learn the entire distribution through adversarial competition.

3. Experimental Results ​

On MNIST, TFD, and other datasets, GAN-generated samples were sharper and more realistic than VAE and other explicit models at the time; the handwritten digits shown in the paper were almost indistinguishable from real ones. But the original GAN also exposed deep problems: training is extremely unstable—if the discriminator is too strong, gradients vanish; too weak, the generator doesn't learn; mode collapse—the generator learns only a few types of samples to fool the discriminator; convergence has no guarantees.

4. Implications for Today ​

  • Implicit modeling changed the entire generation field: as long as a differentiable "real/fake judgment" exists, distributions can be indirectly learned. This thinking was inherited by countless works including conditional GAN, StyleGAN, and domain adaptation.
  • Training instability is its greatest legacy: GAN's adversarial training is far harder than minimizing a single loss. This very nagging problem left enormous room for the later, more stable diffusion models.
  • Adversarial thinking spilled over: adversarial examples (tiny perturbations that fool classifiers), adversarial training, GAN data augmentation—"letting two models make each other suffer" became a universal tool.
  • Best read side-by-side with DDPM: one uses game theory, one uses denoising; one unstable but once the highest quality, one stable and later surpassing in quality. Two philosophies for the same problem, an excellent lens for understanding generative AI.

GAN's methodological lesson for us

The GAN paper created a field, but its own experimental section also truthfully reported "generation results are unstable, requiring careful tuning." A paper's value isn't whether it's perfect, but whether it opens a path that enough people are willing to patch. The subsequent WGAN, DCGAN, and StyleGAN each fixed GAN's stability—this is evidence of "a good paper opens the door."

8. DDPM: Diffusion Models Reshuffle the Generation Table (Ho, 2020) ​

1. Problem Background ​

In 2020, image generation was at a crossroads: GAN had the highest quality but unstable training and mode collapse; VAE was stable but blurry, lacking sharp details. Could there be a method that was stable to train, fully covered modes, and had high generation quality? Ho, Jain, and Abbeel took the diffusion idea (inspired by non-equilibrium thermodynamics) proposed by Sohl-Dickstein et al. in 2015 and re-landed it with an engineering recipe, providing the answer—Denoising Diffusion Probabilistic Models (DDPM).

2. Core Mechanism: Noise-Adding → Noise-Removing ​

Diffusion models are a "pollute first, clean up later" two-stage process:

Forward process (noise-adding, no learning needed):
  x₀ (real image) ──add a bit of noise──▶ x₁ ──▶ x₂ ──▶ ⋯ ──▶ x_T (pure noise)
  Each step: x_t = √(1−β_t)·x_{t−1} + √β_t·ε , ε ~ N(0, I)

Reverse process (noise-removing, learned by neural network):
  Training goal: given noisy image x_t and timestep t, predict the added noise ε
  Generation:      starting from pure noise x_T, denoise step by step, T steps to recover x₀

Three key engineering recipes that took it from "theoretically feasible" to "actually generates":

  • Simplified training objective: reduce the complex variational lower bound (ELBO) to just optimizing the mean squared error of "predicting noise"—the supervision signal is extremely simple and clean.
  • U-Net backbone + timestep embedding: use a U-Net with skip connections as the denoising network, embedding the timestep t as a condition (the network needs to know "how much noise to remove at this point").
  • Unified view with VAE: diffusion models essentially are a special hierarchical VAE—each layer "scatters the input into noise" then learns to recover it, just scattering thoroughly enough.

3. Experimental Results ​

On unconditional CIFAR-10, DDPM reached an FID of 9.46, comparable to or better than the best GANs at the time; yet its training process was stable throughout, with no mode collapse, and generation diversity far exceeded the GAN family. The paper simultaneously proved via sampling quality evaluation (FID/IS, etc.) that this simple route of "random noise-adding + learning to denoise" not only works but wins decisively.

4. Implications for Today ​

  • Stability and quality together: DDPM's fundamental reason for winning is—modeling generation as the simple regression problem of "predicting noise," replacing unstable adversarial games with stable supervised learning. Stability is the key to surpassing GAN and the prerequisite for large-scale engineering.
  • A sample of paradigm shift: when adversarial training was stuck at its bottleneck, a more simple "denoising autoencoder" overtook via a backroad. This reminds us: the frontier isn't won by only one way; sometimes the answer lies in more fundamental math.
  • Subsequent chain reactions: DDIM (faster sampling), Classifier-Free Guidance (controllable generation), Latent Diffusion (backbone of Stable Diffusion) stacked layer upon layer, pushing diffusion models to the peak of 2022+ image/video/audio generation. See Diffusion Model Case Study.
  • Methodology: introducing old ideas from physics/math (thermodynamic diffusion, Langevin dynamics) into ML is a shortcut to discovering new models—diffusion isn't the first example (residuals and attention also have predecessors), and it won't be the last.

Distinguishing GAN from diffusion in one sentence

GAN uses a discriminator to force the generator to "look real"; diffusion uses T steps of small noise to force the denoiser to "step by step restore." The former seeks equilibrium in a game, the latter seeks stability in simple regression. Understanding this difference is understanding why image generation as a whole shifted to diffusion in 2022.

9. AlphaGo: MCTS + Neural Networks, Perfectly Combined (Silver, 2016) ​

1. Problem Background ​

Go is widely acknowledged as AI's hardest-to-conquer board game: the 19×19 board has ~10¹⁷⁰ legal positions (compared to chess's ~10⁴⁷), far beyond the limits of brute-force search; and position evaluation is extremely hard, with no simple heuristic function. Before 2016, the strongest Go AI reached only amateur level. DeepMind's AlphaGo fused deep learning + reinforcement learning + Monte Carlo Tree Search (MCTS) into one.

2. Core Mechanism: Intuition + Deliberation ​

        ┌────────────────────────────────────────────┐
        │            Training Phase (offline, costly)  │
        │  Policy Network: supervised learning from human games → self-play RL optimization  │
        │  Value Network: regression of win rate from self-play data                      │
        └────────────────────────────────────────────┘
                              │ Provides "intuition"
                              ▼
        ┌────────────────────────────────────────────┐
        │            Search Phase (online, fast)       │
        │  MCTS: policy network guides "where to look" (narrowing the branching) │
        │        value network evaluates "how good the position is" (replacing random rollouts) │
        └────────────────────────────────────────────┘
  • Policy Network: inputs the board position, outputs the probability of the next move. First supervised pretrained on human pro game records, then further optimized via self-play + policy gradient RL. It compresses the branching factor from ~250 to dozens—search only unfolds around "promising" branches.
  • Value Network: evaluates the win rate of a given position, replacing the costly random rollouts (rollout) in traditional MCTS, allowing the search to make deeper, more accurate judgments with less compute.
  • MCTS integration: during inference, the policy network handles "which line to explore," the value network handles "what's the outcome of this line"—the two correct each other in tree search—the classic combination of "learning + search." See the RL applications for the RL loop.

3. Experimental Results ​

  • October 2015, AlphaGo defeated European Go champion Fan Hui 5:0—AI defeating a pro player for the first time.
  • March 2016, defeated world champion Lee Sedol 4:1. The 37th move (game 2), the "5th line move," was publicly recognized by pro players as a creative move beyond human game records—AI first demonstrated a chess style outside of humanity in a game that "relies on intuition and beauty."
  • Distributed training version used ~1202 CPUs + 176 GPUs; the single-machine version also defeated Fan Hui.

4. Implications for Today ​

  • Division of labor between learning and search: learning provides "intuition" (where to think, how good a position is), search provides "deliberation" (verifying intuition to the end). This combination route continued to evolve after AlphaGo: AlphaZero eliminated human game records, surpassing human level with pure self-play; today's large model "chain-of-thought + search enhancement" (test-time compute) is exactly the same thinking.
  • Supervised learning to be human, RL to win: AlphaGo used supervised pretraining to get started fast, then RL to surpass humans—the design of training objectives determines the ceiling more than network architecture.
  • Milestone significance: this was the first time deep learning publicly defeated humans in a "pinnacle of human intellect" board game, permanently changing public perception of AI. The full historical narrative is in Evolutionary History.

Echo with the Perceptron

58 years separated AlphaGo from the Perceptron, but they share the same basic belief: structure comes from neuroscience-inspired insights (deep networks modeling the brain's hierarchical processing), capability comes from data-driven learning. The only difference: the Perceptron learned a straight line, AlphaGo learned the entire strategy and value of Go. What grew over sixty years wasn't the belief, but data, compute, and network depth.

10. GPT-3: The Scaling Law Arrives (Brown, 2020) ​

1. Problem Background ​

BERT proved "pretraining + fine-tuning" works, but fine-tuning has a real bottleneck: every new task requires collecting labeled data and retraining, which is costly and hard to scale. GPT-3 asked a more radical question: what if we skip fine-tuning entirely—give the model a few examples (or even zero), and it just does the work? The paper title itself is a manifesto: Language Models are Few-Shot Learners—language models are "few-shot learners."

2. Core Mechanism: Scale + In-Context Learning ​

  • Scale: 175 billion parameters, two orders of magnitude larger than the previous GPT-2; training data: ~45TB cleaned web text (Common Crawl, WebText, etc.). This was the world's largest neural network at the time.
  • In-Context Learning: put task examples directly into the input prompt (prompt), the model "learns on sight," without updating any parameters. E.g., for translation, give two or three "English → Chinese" examples in the prompt, and the model does it directly.
  • Three evaluation settings: zero-shot, one-shot, and few-shot. Results almost always improved monotonically with the number of examples—examples in context are "online fine-tuning."
  • Empirical scaling laws: training loss and downstream task capability improve steadily by power law with parameters, data, and compute—"more compute + more data + bigger model" translates directly and predictably into stronger capability.
Pretraining (one-time, astronomically expensive):
  45TB text ──▶ 175B parameter model (learning "statistical laws of the world")

Inference (learn-on-sight, zero-cost adaptation):
  prompt = "Translate to English: cat → cat; dog → dog; bird →" ──▶ "bird"
  ↑ No gradient update, just a few examples, model can complete the new task

3. Experimental Results ​

Across 20+ tasks including translation, QA, closed-book knowledge filling, arithmetic, and news generation, GPT-3's few-shot performance was comparable to or better than at the time specially fine-tuned SOTA (e.g., TriviaQA closed-book QA, arithmetic reasoning approaching or exceeding fine-tuned models). The paper also truthfully reported the dark side of large models: biases in training data are amplified, generated content may fabricate facts (hallucination), and diminishing returns at large scale. This paper both declared the power of scale and foreshadowed its risks.

4. Implications for Today ​

  • The scaling law is the hardest faith of the 2020s: it directly birthed ChatGPT (2022), GPT-4 (2023), and the global large model arms race. Understanding it is understanding why every company is frantically stacking GPUs and buying data.
  • Another paradigm shift: from "pretraining + fine-tuning" to "pretraining + prompting/alignment"—human-model interaction changed from "writing code for fine-tuning" to "conversing." BERT established the pretraining paradigm; GPT-3 simplified the interaction method to natural language.
  • Capability hides in data: the performance of few-shot learning suggests that many "capabilities" may already be lying dormant in the pretraining data, just needing the right trigger—this changes our understanding of "what the model has learned." See Large Language Models.
  • Scale isn't a magic bullet: bias, hallucination, data scarcity, and physical compute limits mean that "just stacking scale" growth will eventually slow. Today's test-time compute, synthetic data, and algorithmic efficiency optimization are all responses to the limitations of scaling laws. Frontier dynamics are at Frontier Progress.

The two sides of scaling laws

Scaling laws tell us "more compute produces miracles," but also remind us of their boundaries: the bigger the model, the more harshly bias in training data is amplified; total data is finite, corpus will eventually run out; a single GPU's compute growth isn't linear. Treating scaling laws as engineering convictions rather than universal dogmas is the balanced mindset you should bring when reading this paper.

11. The Right Posture for Reading Papers: The Three-Pass Method ​

Deep reading isn't "reading word by word from start to finish"—it's reading in rhythm, three passes, each with different goals and outputs. This method comes from S. Keshav's classic paper How to Read a Paper:

PassTime BudgetWhat to ReadOutput
Pass 1 · Bird's-Eye View5–10 minTitle, abstract, introduction, conclusion, all figure captionsAnswer: what problem does it solve? Core method in one sentence? How well does it work?
Pass 2 · Structure~1 hourFull text body, figures, formulas at a high level, method flowCan draw the method flow diagram, can retell key experiments and ablations
Pass 3 · Deep ReadSeveral hoursLine-by-line derivations, reproduce experiments, critical reviewCan point out limitations, propose improvements, judge "what if in my scenario"

Specific operational advice:

  1. Answer only three questions in Pass 1: what pain point does it solve? why does it claim to be a new contribution? are its conclusions credible? If you can't answer, re-read the abstract—if you can answer in Pass 1, congratulations, you've graduated from "skim-reading" this paper.
  2. Draw figures by hand in Pass 2: draw methods as block diagrams, organize ablation experiments into tables. "Being able to draw it" and "being able to understand it" are the same thing; if you can't draw it, you don't understand yet. This step is worth practicing on the nine papers in this article one by one.
  3. Bring a critical eye in Pass 3: are baselines fair? were hyperparameters tuned? do conclusions hold on out-of-distribution data? then search for follow-up papers, see how this paper was criticized and surpassed—"how it was surpassed" is often more informative than the paper itself.
  4. Note only four things: one-sentence mechanism, key numbers (with conditions), the author's failures and limitations, and relevance to your scenario. Refer to the note template in the site's Reading Paths, and look up terms in the Glossary as you go.

A key judgment: not every paper deserves a Pass 3. Many papers are enough after halfway through Pass 2—when you can predict what the next paragraph says, stop, and save the time for the next article or a reproduction experiment.

Three-pass method in practice (using ResNet as demo)

Pass 1: abstract + conclusion—"residual network with identity shortcuts solves deep degradation, 152 layers exceed human level." Pass 2: read methods + experiments—draw the residual block block diagram, organize evidence for "degradation is not overfitting," compare error tables for 34/50/101/152 layers. Pass 3: derive why gradients propagate back along shortcuts quickly, question "was BN's role counted in the residual gain," then search "why do residual connections work" follow-up analysis papers. Only after three passes does the paper truly belong to you.

12. The Common Patterns Behind These Nine Papers ​

These nine papers span 62 years, covering five fields—perception, language, game, generation—but together, the patterns are strikingly consistent:

Pattern 1: good papers all come from solving specific pain points, not chasing new concepts.

  • The Perceptron solved "can machines learn to judge?"; AlexNet solved "deep CNNs can't train"; ResNet solved "deeper is worse"; Transformer solved "RNN can't parallelize + long-range forgetting"; BERT solved "language models are unidirectional"; GAN solved "generative models can't produce sharp samples"; DDPM solved "GAN training is unstable"; AlphaGo solved "Go's search branching factor is too large"; GPT-3 solved "every task needs labeled data for fine-tuning."
  • Every pain point was a specific obstacle blocking the road at the time, all "important enough to need no explanation" in hindsight. Pain points define problems; problems define innovation.

Pattern 2: core mechanisms are simple enough to explain in one sentence.

  • Residual connection: replace F(x) with F(x) + x; attention: softmax(QKᵀ/√d)·V; Dropout: randomly drop half the neurons; diffusion: add noise then learn to denoise; adversarial: let two networks scam each other. Core mechanisms that change fields are often explainable in one sentence. Complexity is in the implementation, not the concept. If a paper's core mechanism takes three pages to explain, it's probably not that level of contribution.

Pattern 3: all use experiments to prove "the mechanism is what's working," not just giving results.

  • ResNet uses "training error is also higher" to rule out overfitting, locating degradation at the optimization layer; Transformer proves architectural advantage with BLEU + training cost dual metrics; BERT validates components one by one across 11 benchmarks + ablations (what happens if NSP is removed?); AlexNet reports "without Dropout, severe overfitting occurs." Good papers don't just give results—they give experimental designs that "prove the results are caused by the mechanism." This is the part you should grab most when reading.

Pattern 4: breakthroughs almost always come from "cross-disciplinary borrowing."

  • The Perceptron borrowed from neuroscience, residuals borrowed from signal processing's "differential/predictive correction," attention borrowed from information retrieval's "query-key-value," GAN borrowed game theory's Nash equilibrium, diffusion borrowed thermodynamics' entropy increase and Langevin dynamics, AlphaGo borrowed game tree search, GPT-3 borrowed power laws from statistical physics. Standing on the shoulders of other disciplines is the shortest path to reducing innovation cost. To track the frontier, first see what hammer it borrowed from whom.

Pattern 5: each opens a door, not closes one.

  • ResNet made "deeper" possible, Transformer made "bigger" possible, DDPM made "stable generation" possible, GPT-3 made "scale is capability" a testable hypothesis—none closed a field; they made fields explode. A simple standard for judging a paper's value: does it make follow-up work easier, cheaper, and more likely? Rereading these nine papers against this standard, they all pass firmly.

One-sentence summary

Reading papers is less about learning knowledge and more about learning how to discover and solve problems. The shared methodology of these nine is: find a specific, measurable pain point → give a mechanism that's simple and interpretable → use ablation experiments to prove the causal chain between mechanism and effect. If you can reproduce this methodology, you've got the skeleton of independent research. The rest is walking every coordinate on the Paper Map.

Further Reading ​

References ​

All below are real, publicly accessible resources, with direct links to the original papers: