Skip to content

Paper Map

Quick overview Lay out deep learning's key papers across six decades into a single map organized by theme and timeline: foundational era, convolution era, sequence modeling, attention & large models, generative models, efficiency & optimization, and scientific applications. Each paper gets a one-sentence contribution, with a comprehensive table (year / paper / theme / contribution / linked pages).

Paper Map ​

One-sentence definition: The Paper Map is a panoramic view that lays out deep learning's key papers across history, organized by theme and timeline — it answers "where did this field come from, what is the main thread, where is it heading next." It's a positioning tool before close-reading any paper (the accessible version of the timeline is at A Brief History of Deep Learning).

I. How to Read This Map ​

  • The horizontal axis is time: every paper stands on the shoulders of predecessors. Reading in chronological order reveals "how problems were progressively forced out";
  • The vertical axis is theme: five main threads — vision, sequence, generative, efficiency, and scientific applications — are relatively independent yet intertwined. For example, the Transformer started as a sequence model, later dominated vision and generation;
  • Use in combination: the map only provides "one-sentence contributions"; for deep reading, go to Classic Papers Deep Dive; for latest movements, jump to Frontier Advances; to order reading by goal, use Reading Paths.

One Sentence to Understand Paper History

The history of deep learning is the story of capabilities being continuously amplified: single-layer perceptron (1958) → multi-layer backpropagation (1986) → large-scale CNN (2012) → long-sequence models (2014-2017) → pretrained large models (2018-2020) → generation & alignment (2020-). Each step forward relied on the combined force of "more data + more compute + smarter inductive bias."

II. Foundational Era: Neurons, Perceptrons, and Backpropagation (1943-1989) ​

  • 1943 · McCulloch & Pitts: Defined the "formal neuron" in mathematical language, abstracting brain computation into simulatable logical units — the starting point of neural network theory.
  • 1958 · Rosenblatt Perceptron: The first "learnable" model, automatically learning weights from samples using a simple perceptron rule, igniting the first wave of neural network enthusiasm. It confirmed a simple yet profound mechanism: cognition can be decomposed into learnable weights (the perceptron's basic structure is at Neural Network Fundamentals).
  • 1969 · Minsky & Papert's Perceptrons: Rigorously proved that a single-layer perceptron can't learn XOR, doused the enthusiasm of the first wave, and plunged the entire field into winter — this lesson still reminds us that "model expressive power has boundaries."
  • 1986 · Rumelhart, Hinton, Williams Backpropagation: Made multi-layer networks trainable using the chain rule, turning "depth" from a theoretical parameter into actual capability. Its mechanism (forward pass → loss → gradient backpropagation) is the foundation of all training today; derivations are at Backpropagation and Automatic Differentiation.
  • 1998 · LeCun LeNet-5: First end-to-end combination of "convolution + pooling + fully connected + backpropagation" for handwritten digit recognition, deployed in banks for years — the CNN skeleton was thus set; see CNN and Computer Vision for anatomy.

III. Convolution Era: Depth, Width, and Training Stability (2012-2020) ​

  • 2012 · AlexNet: The first year of deep learning. ReLU alleviated gradient vanishing, Dropout prevented overfitting, multi-GPU scaled up networks, and brought top-5 error rates from 26% down to 15.3% on ImageNet. Its significance wasn't architectural innovation but validating the "deep + big + strong compute" path.
  • 2015 · VGG / GoogLeNet: VGG pushed networks to 16-19 layers using uniform small conv kernels (3×3), proving "neatness beats fancy"; GoogLeNet's Inception widened the view with multi-scale parallelism.
  • 2015 · BatchNorm: Normalization per mini-batch turned deep network training from "walking a tightrope" to "walking a broad highway," giving deep learning the courage to use large learning rates. Trade-offs with LayerNorm are at Initialization and Normalization.
  • 2016 · ResNet: Residual connections made 152 layers possible, and the ILSVRC 2015 champion. It turned the "deeper is worse" degradation problem into "adding a layer can only help" — a mechanism inherited by all subsequent architectures including Transformer.
  • 2020 · ViT: Cutting images into patches as tokens, proving "pure Transformer can beat CNN" — the inductive bias of convolution was replaced by large-scale pre-training, and vision and language architectures began to unify (see Multimodal Models).

IV. Sequence Modeling: From LSTM to Attention (1997-2017) ​

  • 1997 · LSTM (Hochreiter & Schmidhuber): Gating mechanisms (forget gate, input gate, output gate) solved the gradient problem for long-range dependencies, dominating sequence modeling for nearly two decades, and still playing roles in speech and streaming tasks today. Full details at RNN and Sequence Modeling.
  • 2013 · Word2Vec: Turning words into dense vectors, where relationships between words became vector operations ("king - man + woman ≈ queen") — the first large-scale validation of the "representation is everything" idea.
  • 2014 · Seq2Seq (Sutskever, Vinyals, Le): The Encoder-Decoder framework allowed neural networks to handle "variable-length to variable-length" mappings (machine translation) for the first time, paving the way for generative models.
  • 2015 · Bahdanau Attention: Letting the decoder "look back" at different parts of the input at each step — the attention mechanism was born, directly planting the seed for Transformer. Concept details at Attention Mechanism.

V. Attention and Large Models: After Transformer (2017-) ​

  • 2017 · Attention Is All You Need: Abandoning recurrence, modeling sequences purely through parallel self-attention. Its triple advantage in complexity, parallelism, and scalability made it the de facto standard of the large model era. Anatomy at Transformer Architecture.
  • 2018-2019 · ELMo / GPT / BERT: Turning "pre-training + fine-tuning" into the mainstream paradigm. BERT's bidirectional masked language model swept NLP leaderboards, while GPT stuck to the autoregressive one-way route — the two-route debate was ultimately won by GPT's "generation unification."
  • 2019 · GPT-2 / T5: GPT-2 proved "bigger is stronger"; T5 unified all NLP tasks into "text-to-text."
  • 2020 · GPT-3: 175B parameters + in-context learning (few-shot prompting), first demonstrating model capability of "no fine-tuning needed, just say a few words." Large-scale scaling laws at Large Language Models (LLM) and Representation Learning and Pre-training.
  • 2022-2023 · InstructGPT / Chinchilla / LLaMA / GPT-4: RLHF made models "speak human"; Chinchilla answered "how much data pairs with how many parameters"; LLaMA's open source redefined the research ecosystem; GPT-4 moved toward multimodality and tools.

VI. Generative Models: From Latent Variables to Diffusion (2014-) ​

  • 2014 · VAE: Learning data's latent distribution through "encode → sample → decode," the "Bayesian school" cornerstone of generative models. See VAE and GAN.
  • 2014 · GAN: Zero-sum game between generator and discriminator; generated images once "passed for real" (StyleGAN faces), but unstable training and mode collapse are lifelong pains.
  • 2020 · DDPM: Step-by-step forward noising, step-by-step reverse denoising — stable training and good diversity, officially starting the diffusion era. Mechanism details at Diffusion Models and Generative AI.
  • 2022 · Stable Diffusion: Putting diffusion into latent space + text conditioning, making "text-to-image" runnable on consumer-grade GPUs.
  • 2024 · Sora: Extending diffusion to video, demonstrating consistent modeling of the physical world. The generative model family tree at Generative Models.

VII. Efficiency and Optimization: Making Large Models "Usable" (2015-) ​

  • 2015 · Adam: Momentum + adaptive learning rate + bias correction in one, becoming the default optimizer. See Optimization and Gradient Descent.
  • 2015 · Knowledge Distillation (Hinton): Small models learning from large models' soft labels, the origin of model compression and the "teacher-student" paradigm.
  • 2017-2021 · MoE (Shazeer 2017 / Switch Transformer 2021): Sparse expert routing — "activate a subset of parameters to do all the work." Large models like DeepSeek-V3 and Mixtral use it to reduce training costs by an order of magnitude.
  • 2022-2023 · FlashAttention / GPTQ / AWQ / Speculative Decoding: Attention I/O optimization, 4-bit quantization, inference acceleration — letting billion-parameter models run on a single GPU or even mobile devices. Full engineering panorama at MLOps and Model Deployment.

VIII. Scientific Applications: Deep Learning Changing Science (2016-) ​

  • 2016 · AlphaGo: Deep reinforcement learning defeating top human players for the first time. See Deep RL Applications.
  • 2021 · AlphaFold2: Pushing protein structure prediction accuracy from "impossible" to experimental level (CASP14 score exceeded 90), predicting nearly all known protein structures within two years.
  • 2024 · AlphaFold3 / AlphaGeometry: Expanding from "predicting structures" to "protein-ligand complexes," and theorem proving at the Math Olympiad level — "AI for Science" went from slogan to daily practice.

IX. Comprehensive Table ​

The following summarizes all the above papers by [year / paper / theme / one-sentence contribution / linked page]. Linked pages are plain text references; detailed links are in each section:

YearPaperThemeOne-Sentence ContributionLinked Page
1943McCulloch & PittsFoundationalMathematically defined the neuron; starting point of NN theoryNeural Network Fundamentals
1958Rosenblatt PerceptronFoundationalFirst learnable model; ignited the first waveNeural Network Fundamentals
1969Minsky & PapertFoundationalProved single-layer perceptron limits; field entered winterNeural Network Fundamentals
1986Rumelhart BackpropagationFoundationalChain rule made multi-layer networks trainableBackpropagation and Automatic Differentiation
1998LeNet-5ConvolutionCNN skeleton of conv+pooling+fully connected setCNN and Computer Vision
2012AlexNetConvolutionFirst validation of deep+big data+strong compute pathCNN and Computer Vision
2015VGGConvolutionUniform small conv kernels; neaten and deepen networksCNN and Computer Vision
2015GoogLeNetConvolutionInception multi-scale parallelism; widen networksCNN and Computer Vision
2015BatchNormConvolutionNormalization stabilizes deep training; use large LRInitialization and Normalization
2016ResNetConvolutionResidual connections made ultra-deep networks trainableCNN and Computer Vision
2020ViTConvolutionImage patchification; Transformer enters visionMultimodal Models
1997LSTMSequenceGating solves long-range dependency; dominated SM for 20 yrsRNN and Sequence Modeling
2013Word2VecSequenceWords become dense vectors; representation learning milestoneRepresentation Learning and Pre-training
2014Seq2SeqSequenceVariable-to-variable mapping frameworkRNN and Sequence Modeling
2015Bahdanau AttentionSequenceAttention mechanism born; seed of TransformerAttention Mechanism
2017Attention Is All You NeedLarge ModelsPure attention parallel modeling; opened the large model eraTransformer Architecture
2018ELMo / GPT-1Large ModelsPre-training + fine-tuning paradigm establishedRepresentation Learning and Pre-training
2019BERTLarge ModelsBidirectional masked pre-training swept NLP leaderboardsLarge Language Models (LLM)
2019GPT-2 / T5Large ModelsBigger is stronger / all tasks unified as text-to-textLarge Language Models (LLM)
2020GPT-3Large Models175B parameters; in-context learning goes mainstreamLarge Language Models (LLM)
2022InstructGPT / ChinchillaLarge ModelsRLHF alignment / compute-optimal ratiosLarge Language Models (LLM)
2023LLaMA / GPT-4Large ModelsOpen-source ecosystem / multimodality and tool useMultimodal Models
2014VAEGenerativeLatent variables + reparameterization for learning data distributionVAE and GAN
2014GANGenerativeGenerator-discriminator adversarial gameVAE and GAN
2019StyleGANGenerativeControllable high-quality face generationVAE and GAN
2020DDPMGenerativeNoising-denoising diffusion generation paradigmDiffusion Models and Generative AI
2022Stable DiffusionGenerativeLatent-space diffusion + text conditioning; democratized T2IDiffusion Models and Generative AI
2024SoraGenerativeDiffusion extended to long-form videoDiffusion Models and Generative AI
2015AdamEfficiencyMomentum + adaptive LR default optimizerOptimization and Gradient Descent
2015Knowledge DistillationEfficiencySmall models learn large models' soft labelsOptimization and Gradient Descent
2017Sparsely-Gated MoEEfficiencySparse expert routing; activate subset of parametersLarge Language Models (LLM)
2022FlashAttentionEfficiencyIO-aware attention accelerationLarge Language Models (LLM)
2022-23GPTQ / AWQ / Speculative DecodingEfficiencyQuantization and speculative decoding reduce inference costsMLOps and Model Deployment
2016AlphaGoScientific ApplicationsDeep RL defeats human playersDeep RL Applications
2021AlphaFold2Scientific ApplicationsProtein structure prediction reaches experimental levelDeep RL Applications
2024AlphaFold3 / AlphaGeometryScientific ApplicationsComplex structure / Olympiad-level theorem provingDeep RL Applications

X. Trade-offs ​

  • Map ≠ knowledge: maps help you position quickly, but "one-sentence contributions" obscure many details and failures — for example, ResNet's success was built on VGG's depth lessons, and a one-liner taken out of context distorts the truth.
  • Main thread vs. side threads: this map highlights seven main threads, but side threads like reinforcement learning (beyond AlphaGo: recommendation, control), GNNs, and speech are equally important — dig into the corresponding sections.
  • Limitations of historical perspective: all maps carry a "hindsight bias." Remember: back then, no one knew which paper would become canonical. Many papers overlooked at the time (like early Transformer, early RNN critiques) were later repeatedly rediscovered.
  • Don't memorize the map: a map's value is "positioning," not "memory." When encountering an unfamiliar model, ask "which main thread is it on, what did it inherit, what did it challenge?" — that's enough.

Further Reading ​

References ​