Theme
Paper Map
One-sentence definition: The Paper Map is a panoramic view that lays out deep learning's key papers across history, organized by theme and timeline — it answers "where did this field come from, what is the main thread, where is it heading next." It's a positioning tool before close-reading any paper (the accessible version of the timeline is at A Brief History of Deep Learning).
I. How to Read This Map
- The horizontal axis is time: every paper stands on the shoulders of predecessors. Reading in chronological order reveals "how problems were progressively forced out";
- The vertical axis is theme: five main threads — vision, sequence, generative, efficiency, and scientific applications — are relatively independent yet intertwined. For example, the Transformer started as a sequence model, later dominated vision and generation;
- Use in combination: the map only provides "one-sentence contributions"; for deep reading, go to Classic Papers Deep Dive; for latest movements, jump to Frontier Advances; to order reading by goal, use Reading Paths.
One Sentence to Understand Paper History
The history of deep learning is the story of capabilities being continuously amplified: single-layer perceptron (1958) → multi-layer backpropagation (1986) → large-scale CNN (2012) → long-sequence models (2014-2017) → pretrained large models (2018-2020) → generation & alignment (2020-). Each step forward relied on the combined force of "more data + more compute + smarter inductive bias."
II. Foundational Era: Neurons, Perceptrons, and Backpropagation (1943-1989)
- 1943 · McCulloch & Pitts: Defined the "formal neuron" in mathematical language, abstracting brain computation into simulatable logical units — the starting point of neural network theory.
- 1958 · Rosenblatt Perceptron: The first "learnable" model, automatically learning weights from samples using a simple perceptron rule, igniting the first wave of neural network enthusiasm. It confirmed a simple yet profound mechanism: cognition can be decomposed into learnable weights (the perceptron's basic structure is at Neural Network Fundamentals).
- 1969 · Minsky & Papert's Perceptrons: Rigorously proved that a single-layer perceptron can't learn XOR, doused the enthusiasm of the first wave, and plunged the entire field into winter — this lesson still reminds us that "model expressive power has boundaries."
- 1986 · Rumelhart, Hinton, Williams Backpropagation: Made multi-layer networks trainable using the chain rule, turning "depth" from a theoretical parameter into actual capability. Its mechanism (forward pass → loss → gradient backpropagation) is the foundation of all training today; derivations are at Backpropagation and Automatic Differentiation.
- 1998 · LeCun LeNet-5: First end-to-end combination of "convolution + pooling + fully connected + backpropagation" for handwritten digit recognition, deployed in banks for years — the CNN skeleton was thus set; see CNN and Computer Vision for anatomy.
III. Convolution Era: Depth, Width, and Training Stability (2012-2020)
- 2012 · AlexNet: The first year of deep learning. ReLU alleviated gradient vanishing, Dropout prevented overfitting, multi-GPU scaled up networks, and brought top-5 error rates from 26% down to 15.3% on ImageNet. Its significance wasn't architectural innovation but validating the "deep + big + strong compute" path.
- 2015 · VGG / GoogLeNet: VGG pushed networks to 16-19 layers using uniform small conv kernels (3×3), proving "neatness beats fancy"; GoogLeNet's Inception widened the view with multi-scale parallelism.
- 2015 · BatchNorm: Normalization per mini-batch turned deep network training from "walking a tightrope" to "walking a broad highway," giving deep learning the courage to use large learning rates. Trade-offs with LayerNorm are at Initialization and Normalization.
- 2016 · ResNet: Residual connections made 152 layers possible, and the ILSVRC 2015 champion. It turned the "deeper is worse" degradation problem into "adding a layer can only help" — a mechanism inherited by all subsequent architectures including Transformer.
- 2020 · ViT: Cutting images into patches as tokens, proving "pure Transformer can beat CNN" — the inductive bias of convolution was replaced by large-scale pre-training, and vision and language architectures began to unify (see Multimodal Models).
IV. Sequence Modeling: From LSTM to Attention (1997-2017)
- 1997 · LSTM (Hochreiter & Schmidhuber): Gating mechanisms (forget gate, input gate, output gate) solved the gradient problem for long-range dependencies, dominating sequence modeling for nearly two decades, and still playing roles in speech and streaming tasks today. Full details at RNN and Sequence Modeling.
- 2013 · Word2Vec: Turning words into dense vectors, where relationships between words became vector operations ("king - man + woman ≈ queen") — the first large-scale validation of the "representation is everything" idea.
- 2014 · Seq2Seq (Sutskever, Vinyals, Le): The Encoder-Decoder framework allowed neural networks to handle "variable-length to variable-length" mappings (machine translation) for the first time, paving the way for generative models.
- 2015 · Bahdanau Attention: Letting the decoder "look back" at different parts of the input at each step — the attention mechanism was born, directly planting the seed for Transformer. Concept details at Attention Mechanism.
V. Attention and Large Models: After Transformer (2017-)
- 2017 · Attention Is All You Need: Abandoning recurrence, modeling sequences purely through parallel self-attention. Its triple advantage in complexity, parallelism, and scalability made it the de facto standard of the large model era. Anatomy at Transformer Architecture.
- 2018-2019 · ELMo / GPT / BERT: Turning "pre-training + fine-tuning" into the mainstream paradigm. BERT's bidirectional masked language model swept NLP leaderboards, while GPT stuck to the autoregressive one-way route — the two-route debate was ultimately won by GPT's "generation unification."
- 2019 · GPT-2 / T5: GPT-2 proved "bigger is stronger"; T5 unified all NLP tasks into "text-to-text."
- 2020 · GPT-3: 175B parameters + in-context learning (few-shot prompting), first demonstrating model capability of "no fine-tuning needed, just say a few words." Large-scale scaling laws at Large Language Models (LLM) and Representation Learning and Pre-training.
- 2022-2023 · InstructGPT / Chinchilla / LLaMA / GPT-4: RLHF made models "speak human"; Chinchilla answered "how much data pairs with how many parameters"; LLaMA's open source redefined the research ecosystem; GPT-4 moved toward multimodality and tools.
VI. Generative Models: From Latent Variables to Diffusion (2014-)
- 2014 · VAE: Learning data's latent distribution through "encode → sample → decode," the "Bayesian school" cornerstone of generative models. See VAE and GAN.
- 2014 · GAN: Zero-sum game between generator and discriminator; generated images once "passed for real" (StyleGAN faces), but unstable training and mode collapse are lifelong pains.
- 2020 · DDPM: Step-by-step forward noising, step-by-step reverse denoising — stable training and good diversity, officially starting the diffusion era. Mechanism details at Diffusion Models and Generative AI.
- 2022 · Stable Diffusion: Putting diffusion into latent space + text conditioning, making "text-to-image" runnable on consumer-grade GPUs.
- 2024 · Sora: Extending diffusion to video, demonstrating consistent modeling of the physical world. The generative model family tree at Generative Models.
VII. Efficiency and Optimization: Making Large Models "Usable" (2015-)
- 2015 · Adam: Momentum + adaptive learning rate + bias correction in one, becoming the default optimizer. See Optimization and Gradient Descent.
- 2015 · Knowledge Distillation (Hinton): Small models learning from large models' soft labels, the origin of model compression and the "teacher-student" paradigm.
- 2017-2021 · MoE (Shazeer 2017 / Switch Transformer 2021): Sparse expert routing — "activate a subset of parameters to do all the work." Large models like DeepSeek-V3 and Mixtral use it to reduce training costs by an order of magnitude.
- 2022-2023 · FlashAttention / GPTQ / AWQ / Speculative Decoding: Attention I/O optimization, 4-bit quantization, inference acceleration — letting billion-parameter models run on a single GPU or even mobile devices. Full engineering panorama at MLOps and Model Deployment.
VIII. Scientific Applications: Deep Learning Changing Science (2016-)
- 2016 · AlphaGo: Deep reinforcement learning defeating top human players for the first time. See Deep RL Applications.
- 2021 · AlphaFold2: Pushing protein structure prediction accuracy from "impossible" to experimental level (CASP14 score exceeded 90), predicting nearly all known protein structures within two years.
- 2024 · AlphaFold3 / AlphaGeometry: Expanding from "predicting structures" to "protein-ligand complexes," and theorem proving at the Math Olympiad level — "AI for Science" went from slogan to daily practice.
IX. Comprehensive Table
The following summarizes all the above papers by [year / paper / theme / one-sentence contribution / linked page]. Linked pages are plain text references; detailed links are in each section:
| Year | Paper | Theme | One-Sentence Contribution | Linked Page |
|---|---|---|---|---|
| 1943 | McCulloch & Pitts | Foundational | Mathematically defined the neuron; starting point of NN theory | Neural Network Fundamentals |
| 1958 | Rosenblatt Perceptron | Foundational | First learnable model; ignited the first wave | Neural Network Fundamentals |
| 1969 | Minsky & Papert | Foundational | Proved single-layer perceptron limits; field entered winter | Neural Network Fundamentals |
| 1986 | Rumelhart Backpropagation | Foundational | Chain rule made multi-layer networks trainable | Backpropagation and Automatic Differentiation |
| 1998 | LeNet-5 | Convolution | CNN skeleton of conv+pooling+fully connected set | CNN and Computer Vision |
| 2012 | AlexNet | Convolution | First validation of deep+big data+strong compute path | CNN and Computer Vision |
| 2015 | VGG | Convolution | Uniform small conv kernels; neaten and deepen networks | CNN and Computer Vision |
| 2015 | GoogLeNet | Convolution | Inception multi-scale parallelism; widen networks | CNN and Computer Vision |
| 2015 | BatchNorm | Convolution | Normalization stabilizes deep training; use large LR | Initialization and Normalization |
| 2016 | ResNet | Convolution | Residual connections made ultra-deep networks trainable | CNN and Computer Vision |
| 2020 | ViT | Convolution | Image patchification; Transformer enters vision | Multimodal Models |
| 1997 | LSTM | Sequence | Gating solves long-range dependency; dominated SM for 20 yrs | RNN and Sequence Modeling |
| 2013 | Word2Vec | Sequence | Words become dense vectors; representation learning milestone | Representation Learning and Pre-training |
| 2014 | Seq2Seq | Sequence | Variable-to-variable mapping framework | RNN and Sequence Modeling |
| 2015 | Bahdanau Attention | Sequence | Attention mechanism born; seed of Transformer | Attention Mechanism |
| 2017 | Attention Is All You Need | Large Models | Pure attention parallel modeling; opened the large model era | Transformer Architecture |
| 2018 | ELMo / GPT-1 | Large Models | Pre-training + fine-tuning paradigm established | Representation Learning and Pre-training |
| 2019 | BERT | Large Models | Bidirectional masked pre-training swept NLP leaderboards | Large Language Models (LLM) |
| 2019 | GPT-2 / T5 | Large Models | Bigger is stronger / all tasks unified as text-to-text | Large Language Models (LLM) |
| 2020 | GPT-3 | Large Models | 175B parameters; in-context learning goes mainstream | Large Language Models (LLM) |
| 2022 | InstructGPT / Chinchilla | Large Models | RLHF alignment / compute-optimal ratios | Large Language Models (LLM) |
| 2023 | LLaMA / GPT-4 | Large Models | Open-source ecosystem / multimodality and tool use | Multimodal Models |
| 2014 | VAE | Generative | Latent variables + reparameterization for learning data distribution | VAE and GAN |
| 2014 | GAN | Generative | Generator-discriminator adversarial game | VAE and GAN |
| 2019 | StyleGAN | Generative | Controllable high-quality face generation | VAE and GAN |
| 2020 | DDPM | Generative | Noising-denoising diffusion generation paradigm | Diffusion Models and Generative AI |
| 2022 | Stable Diffusion | Generative | Latent-space diffusion + text conditioning; democratized T2I | Diffusion Models and Generative AI |
| 2024 | Sora | Generative | Diffusion extended to long-form video | Diffusion Models and Generative AI |
| 2015 | Adam | Efficiency | Momentum + adaptive LR default optimizer | Optimization and Gradient Descent |
| 2015 | Knowledge Distillation | Efficiency | Small models learn large models' soft labels | Optimization and Gradient Descent |
| 2017 | Sparsely-Gated MoE | Efficiency | Sparse expert routing; activate subset of parameters | Large Language Models (LLM) |
| 2022 | FlashAttention | Efficiency | IO-aware attention acceleration | Large Language Models (LLM) |
| 2022-23 | GPTQ / AWQ / Speculative Decoding | Efficiency | Quantization and speculative decoding reduce inference costs | MLOps and Model Deployment |
| 2016 | AlphaGo | Scientific Applications | Deep RL defeats human players | Deep RL Applications |
| 2021 | AlphaFold2 | Scientific Applications | Protein structure prediction reaches experimental level | Deep RL Applications |
| 2024 | AlphaFold3 / AlphaGeometry | Scientific Applications | Complex structure / Olympiad-level theorem proving | Deep RL Applications |
X. Trade-offs
- Map ≠ knowledge: maps help you position quickly, but "one-sentence contributions" obscure many details and failures — for example, ResNet's success was built on VGG's depth lessons, and a one-liner taken out of context distorts the truth.
- Main thread vs. side threads: this map highlights seven main threads, but side threads like reinforcement learning (beyond AlphaGo: recommendation, control), GNNs, and speech are equally important — dig into the corresponding sections.
- Limitations of historical perspective: all maps carry a "hindsight bias." Remember: back then, no one knew which paper would become canonical. Many papers overlooked at the time (like early Transformer, early RNN critiques) were later repeatedly rediscovered.
- Don't memorize the map: a map's value is "positioning," not "memory." When encountering an unfamiliar model, ask "which main thread is it on, what did it inherit, what did it challenge?" — that's enough.
Further Reading
- Classic Papers Deep Dive — close reading of the most important papers on the map
- Reading Paths — reading order by goal (introduction / deployment / research)
- Frontier Advances — the "latest page" of the map
- Reading Discipline & FAQ — methodology for reading papers
- A Brief History of Deep Learning — the timeline version of the same history
- What Is Deep Learning — domain positioning outside the map
References
- McCulloch, Pitts. A Logical Calculus of the Ideas Immanent in Nervous Activity (1943)
- Rumelhart, Hinton, Williams. Learning representations by back-propagating errors (Nature 1986)
- LeCun et al. Gradient-Based Learning Applied to Document Recognition (Proceedings of the IEEE 1998)
- Krizhevsky, Sutskever, Hinton. ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012)
- He et al. Deep Residual Learning for Image Recognition (CVPR 2016)
- Hochreiter, Schmidhuber. Long Short-Term Memory (Neural Computation 1997)
- Sutskever, Vinyals, Le. Sequence to Sequence Learning with Neural Networks (NeurIPS 2014)
- Bahdanau, Cho, Bengio. Neural Machine Translation by Jointly Learning to Align and Translate (ICLR 2015)
- Vaswani et al. Attention Is All You Need (NeurIPS 2017)
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019)
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020)
- Goodfellow et al. Generative Adversarial Nets (NeurIPS 2014)
- Ho, Jain, Abbeel. Denoising Diffusion Probabilistic Models (NeurIPS 2020)
- Jumper et al. Highly accurate protein structure prediction with AlphaFold (Nature 2021)
- Silver et al. Mastering the game of Go with deep neural networks and tree search (Nature 2016)