Skip to content

Paper Map

Quick overview Organizing key papers from 60 years of machine learning development into a searchable map grouped by theme, covering seven major themes: ML foundations, deep learning foundations, Transformer, generative models, reinforcement learning, tree models, and representation learning. Includes a paradigm shift timeline, paper index by task, and deep/skim reading routes.

Paper Map ​

In one sentence: This map organizes key papers from 60 years of ML development into a searchable coordinate system by theme and era—from the 1958 Perceptron to the 2022 Latent Diffusion. After reading this map, you'll be able to articulate "what problem each classic paper solved, which prior and follow-up works it connects to, and whether it's worth deep-reading or skimming."

The biggest barrier to reading papers isn't "not understanding one paper"—it's "not knowing which paper to read next." Tutorials teach in order, but the paper world is organized by problems: the answer to the same problem (e.g., "how should machines understand images") was SVM + handcrafted features in the 1990s, AlexNet in 2012, and CLIP in 2021. If you only read chronologically, you'll get lost in the details of "who replaced whom"; if you only read by theme, you'll miss the macro narrative of "why paradigms shift." This map's purpose is to layer theme and time so that in front of any paper, you can answer three questions: which technical line does it belong to? What stage of paradigm evolution does it represent? Which papers form a "must-read chain" with it?

How this map relates to other pages

The papers section of this site has five pages with different roles: Start Here discusses "why read papers," Reading Paths discusses "how to read based on your goal," this page (Paper Map) discusses "which papers across the field are worth reading and how they relate to each other," Classic Paper Deep Dives deep-dives each paper, and Frontier Progress tracks new 2020s trends. In short: the map sets coordinates, deep reading provides depth, and the frontier points direction.

1. How to Read This Map ​

Before unfolding the map, let me clarify three reading conventions:

  1. Numbering convention: Papers after 2000 (since arXiv was founded) have an arXiv number in the table, and full text is freely available at https://arxiv.org/abs/number; earlier classic papers have no arXiv number, marked with "—" in the table, with venue listed in the reference section where you can find original links.
  2. Year convention: Paper years always refer to the first public release (formal conference/journal publication, or arXiv preprint submission). For example, VAE's arXiv preprint was submitted in December 2013, ICLR 2014 publication—the table records it as 2014 (preprint 2013). The labeling rule matters more than the specific number.
  3. Deep-read convention: The map only gives a "one-sentence contribution," answering "why is this paper worth remembering." To fully absorb a paper, jump to Classic Paper Deep Dives; to follow its subsequent development, use Frontier Progress and the retrieval method in the FAQ.

Use all three coordinates together: theme determines which line it belongs to, year determines its position in paradigm evolution, and significance determines how much time to invest. Below is the map overview:

                      Theme dimension (seven technical lines)
                    ┌──────────────────────────────────────┐
   Time dimension ↓  │  ① ML foundations                   │
                    │  ② Deep learning foundations          │
                    │  ③ Attention & Transformer            │
                    │  ④ Generative models                  │
                    │  ⑤ Reinforcement learning             │
                    │  ⑥ Tree models & ensembles            │
                    │  ⑦ Representation & self-supervised   │
                    └──────────────────────────────────────┘

2. Academic Map Grouped by Theme ​

1. ML Foundations: Three Pillars (1958–1997) ​

Modern ML has three roots: the neural network line (perceptron and its descendants), the geometric line (maximum-margin classifiers), and the statistical learning line (boosting and ensembles). Before the 1990s they belonged to different schools; in the 1990s they converged on the common battlefield of "tabular data classification," forming the complete arsenal of the feature engineering era.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
The PerceptronFrank Rosenblatt (Cornell Aeronautical Laboratory)1958— (Psychological Review)The first trainable neural network model, proving "machines can learn rules from samples," igniting the first AI wave
SVM (Support-Vector Networks)Corinna Cortes & Vladimir Vapnik (AT&T)1995— (Machine Learning 20)Maximum margin + kernel trick, producing the strongest classifier of the feature engineering era—the pinnacle of geometry and statistics
AdaBoostYoav Freund & Robert Schapire (UC San Diego)1997— (JCSS 55)Proved that weak learners can be "boosted" into strong learners, laying the dual theoretical and practical foundations of ensemble learning

Why read these three first

The Perceptron tells you about the shift in thinking "from rules to learning"; SVM tells you "what the best solution of the feature engineering era looks like"; AdaBoost tells you "combinations of weak models can become strong." Only reading all three together lets you understand why deep learning after 2012 was a revolution rather than an evolution—it simultaneously replaced both the "handcrafted features" and "shallow models" assumptions. The replacement and supplementation of deep learning to these three lines is covered in Deep Learning Foundations.

2. Deep Learning Foundations: Compute, Architecture, and Training (2012–2015) ​

The four papers in this group each solve one of deep learning's three problems: how to use compute (GPU), how to build networks (depth and residuals), and how to train stably (normalization). After reading all four, you'll have mastered the entire engineering recipe of 2010s deep learning.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
AlexNetKrizhevsky, Sutskever, Hinton (University of Toronto)2012arXiv:1207.0580GPU + ReLU + Dropout training of deep CNN, cliff-edge drop in ImageNet error rate, igniting the deep learning revolution
VGGKaren Simonyan & Andrew Zisserman (University of Oxford)2014arXiv:1409.1556Stacking uniform 3×3 small convolutions, empirically proving "deeper is better," becoming the universal backbone for subsequent vision models
ResNetHe, Zhang, Ren, Sun (Microsoft Research Asia)2015arXiv:1512.03385Residual connections make hundreds-layer networks trainable, ImageNet top-5 error 3.57%, first exceeding human baseline
Batch NormalizationSergey Ioffe & Christian Szegedy (Google)2015arXiv:1502.03167Normalizes each layer's input, greatly stabilizing deep network training, turning "deeper" from a slogan into reality

Remember this group in one sentence

AlexNet proved "it can learn," VGG proved "it needs to be deep," ResNet proved "deep has a solution," BatchNorm proved "training can be stabilized." Together, these four form the complete recipe of deep learning's first industrialization. Interestingly, ResNet and BatchNorm appeared in the same year and mutually reinforced each other—without BatchNorm, a 152-layer ResNet would be hard to train; without residuals, the gains from normalization would be offset by degradation.

3. Attention & Transformer: From Component to Universal Base (2014–2020) ​

The attention mechanism started as an "alignment component" in machine translation, was upgraded by the 2017 Transformer to an entire architecture, pretraining paradigms took over NLP starting in 2018, and GPT-3 in 2020 turned "more compute produces miracles" into an empirically-grounded law. This line is the direct ancestor of today's large language models (LLMs); see Transformer Case Study and Large Language Models for mechanism details.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
Neural Machine Translation by Jointly Learning to Align and TranslateBahdanau, Cho, Bengio (Université de Montréal)2014arXiv:1409.0473Introduced attention to let the decoder dynamically "align" with the source sentence, solving the information bottleneck of long-sentence translation—the origin of the attention concept
Attention Is All You NeedVaswani et al. (Google Brain)2017arXiv:1706.03762Proposed the pure-attention architecture Transformer, removing recurrence and convolution, becoming the shared base for BERT/GPT and all large models
BERTDevlin et al. (Google)2018arXiv:1810.04805Bidirectional Transformer + pretraining/fine-tuning paradigm, swept 11 NLP benchmarks, established the "pretrain then fine-tune" standard workflow
GPT-2Radford et al. (OpenAI)2019arXiv:1902.09599Larger-scale autoregressive language model demonstrates zero-shot transfer potential, making "generation" rather than "fill-in-the-blank" the protagonist for the first time
GPT-3Brown et al. (OpenAI)2020arXiv:2005.14165175 billion parameters + in-context learning, empirically proving "scale itself drives capability leaps," kickstarting the scaling law narrative

A common point of confusion

BERT and GPT follow two different pretraining routes: BERT is a bidirectional encoder, good at "understanding" (classification, extraction, retrieval), while GPT is an autoregressive decoder, good at "generation." The 2020s LLM competition ultimately saw the generation route win—ChatGPT proved that "treating everything as generation" is a more unified paradigm. This divergence and convergence is a key thread for understanding NLP history from 2018 to 2024; read alongside Evolutionary History.

4. Generative Models: From Adversarial to Diffusion (2013–2022) ​

Generative models answer only one question: given a bunch of data, can we learn to "create new data with the same distribution"? In 2014, two routes appeared simultaneously—VAE followed the "explicit probability" route of latent variables + variational inference, GAN followed the "implicit distribution" route of adversarial game; in 2020, diffusion models ended GAN's dominance in image generation with more stable training and higher sample quality; in 2022, Latent Diffusion moved diffusion into latent space, making "text-to-image" a consumer product.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
VAE (Auto-Encoding Variational Bayes)Kingma & Welling (University of Amsterdam)2014arXiv:1312.6114Trains a latent-variable model that can generate new samples using variational lower bound, the starting point for latent space generation and the "reparameterization trick"
GAN (Generative Adversarial Nets)Goodfellow et al. (Université de Montréal)2014arXiv:1406.2661Zero-sum game between generator and discriminator, "creating something from nothing" via noise, igniting adversarial generation research
DDPM (Denoising Diffusion Probabilistic Models)Ho, Jain, Abbeel (UC Berkeley)2020arXiv:2006.11239Uses a stepwise denoising Markov chain for generation—stable training, high sample quality, laying the foundation for diffusion-based generation
Latent Diffusion ModelsRombach et al. (LMU Munich et al.)2022arXiv:2112.10752Diffusion in latent space + text-condition injection, efficient and controllable training, becoming the foundation for Stable Diffusion and other image generation tools

How the three routes relate

If VAE is "compress then reconstruct" (with explicit latent variables and probabilistic interpretation), GAN is "counterfeiters and authenticators co-evolving" (no explicit probability but surprisingly sharp), diffusion is "destroy then restore" (gradually denoise from pure noise). The 2020s landscape: diffusion won in images, but VAE's latent space thinking and GAN's adversarial thinking still live in many new methods—Latent Diffusion itself is "VAE's latent space + diffusion's generator." Frontier evolution is covered in Frontier Progress.

5. Reinforcement Learning: From Games to Alignment (2015–2022) ​

What makes RL unique: supervised learning has labels, RL only has "delayed reward." This group of papers showcases RL's two leaps—the first from "handcrafted features" to "deep networks reading raw inputs" (DQN), the second from "playing games" to "aligning large models" (RLHF). The latter directly birthed ChatGPT, and its significance goes far beyond gaming.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
DQN (Human-level control through deep RL)Mnih et al. (DeepMind)2015— (Nature 518)Deep Q-network learned to play 49 Atari games from raw pixels, many surpassing humans, founding deep reinforcement learning
AlphaGoSilver et al. (DeepMind)2016— (Nature 529)Deep policy/value networks + Monte Carlo Tree Search defeated a Go world champion—the classic combination of search and learning
AlphaZeroSilver et al. (DeepMind)2017arXiv:1712.01815Mastered Go/Chess/Shogi from scratch without human game records, pure self-play—the benchmark for general game intelligence
PPOSchulman et al. (OpenAI)2017arXiv:1707.06347Clipped-objective policy gradient algorithm—stable, easy to implement, the de facto standard for policy optimization
RLHF / InstructGPTOuyang et al. (OpenAI)2022arXiv:2203.02155Human feedback + RL (with PPO as the engine) aligns large models to user intent—the core technology of ChatGPT

On the positioning of RLHF

Strictly speaking, InstructGPT (the representative RLHF paper) is a crossover product of "large models + RL": pretraining handles "ability to speak," RLHF handles "ability to speak properly." It belongs to both [Attention & Transformer group] and this group's narrative—this is exactly the map's value: a paper can span multiple themes, and when reading, you should hang it on two lines simultaneously. Tracking the "alignment" frontier is covered in Frontier Progress.

6. Tree Models & Ensembles: The Evergreen of Tabular Data (2001–2017) ​

This group might be the "most easily overshadowed by deep learning" important map. The fact is: on structured tabular data (financial risk control, ad click prediction, sales forecasting), gradient boosting trees remain the kings of production—Kaggle competition winners heavily rely on XGBoost/LightGBM/CatBoost, and deep learning often can't beat them.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
Random ForestsLeo Breiman (UC Berkeley)2001— (Machine Learning 45)Bootstrap sampling + random feature subset ensemble of decision trees, anti-overfitting, nearly zero-tuning, still a universal baseline
XGBoostTianqi Chen & Carlos Guestrin (University of Washington)2016arXiv:1603.02754Second-order gradient + regularized GBDT engineering implementation, the dominant tool in Kaggle and industry from 2015–2020
LightGBMKe et al. (Microsoft)2017arXiv:1706.08374Histogram binning + leaf-wise growth strategy, fast training, low memory, the standard for massive tabular data
CatBoostProkhorenkova et al. (Yandex)2017arXiv:1706.09516Symmetric trees + ordered boosting, native handling of categorical features, anti-overfitting, minimal tuning burden

Three reasons to learn tree models

First, they are the workhorses of production—most companies' core revenue models are tree models, not large models; second, they are the best textbook for understanding "bias-variance tradeoff," "ensembles," and "feature importance," with the most intuitive concepts; third, they are the control group for deep learning—knowing that tree models are stronger on tabular data prevents you from "blindly using deep learning." The site's core knowledge on evaluation and regularization often uses tree models as examples.

7. Representation & Self-Supervised: Letting Models Find Their Own Supervision (2013–2021) ​

This group answers "what to do without labels." Self-supervised learning = constructing supervisory signals from the data itself: Word2Vec makes "word co-occurrence in context" the supervision, CLIP makes "images paired with their captions" the supervision, SimCLR makes "two augmented views of the same image" the supervision. Together, they push ML from "label-driven" to "data-driven," forming the philosophical foundation for why large model "pretraining" works.

PaperAuthor/InstitutionYearVenue/arXivOne-Sentence Contribution
Word2VecMikolov et al. (Google)2013arXiv:1301.3781Learns distributed word representations via neural networks, demonstrating "king − man + woman ≈ queen" semantic arithmetic—the origin of NLP representation learning
SimCLRChen et al. (Google)2020arXiv:2002.05709Contrastive learning framework: unlabeled images can learn strong visual representations, refreshing self-supervised vision performance
CLIPRadford et al. (OpenAI)2021arXiv:2103.00020Contrastive learning on 400M image-text pairs, giving vision models zero-shot transfer ability to "look at an image and say what it is in words," connecting the text and image modalities

Why self-supervised learning is the dark line of the 2020s

The ceiling of supervised learning is labels: human annotation is expensive, error-prone, and incomplete. Self-supervised turns tasks that are forever free and unlimited—"what's the next sentence?" "is this input noisy or the same image?"—into pretraining objectives. Models first learn language syntax and image structure, then fine-tune with a few labels. BERT's "masked word fill," GPT's "predict the next word," CLIP's "image-text matching"—they are all different instances of the same core idea. Term quick-reference is at the Glossary.

3. Paradigm Shift Timeline: Rules → Features → Representations → Scale ​

Projecting the seven themes from Section 2 onto a timeline reveals a clear "paradigm shift": the field's answer to "where does intelligence come from?" went through four major revisions.

   Rules Era            Feature Engineering Era    Representation Learning Era   Scale Era
   1950s–1980s          1990s–2000s              2012–2018                     2018–present
───────────────────────┬─────────────────────────┬───────────────────────────────┬────────────────────
 Intelligence = Rules │ Intelligence = Features  │ Intelligence = Architecture  │ Intelligence = Scale + Alignment
 Expert systems, logic│ + tuning                  + Data                        Pretrained large models
 Knowledge bases,     │ Handcrafted features +    CNN / RNN / Transformer       BERT / GPT / Diffusion
 search               │ shallow models            Networks auto-learn features  Data + compute + alignment
                      │ Random forests            Networks learn features       
                      │ Humans write features     Humans write features (prompt design)

Key paper nodes on each era (year of first public release):

1958 ─ Perceptron                                     (the first "learn from data" voice in the rules era)
1986 ─ Backpropagation (Rumelhart, Hinton, Williams)
1995 ─ SVM                                           (the pinnacle of feature engineering era)
1997 ─ AdaBoost
2001 ─ Random Forests
2006 ─ Deep Belief Networks (Hinton et al.)            (the harbinger of deep learning revival)
2012 ─ AlexNet ◀────────────────────────────          The watershed of the representation learning era
2014 ─ Seq2Seq / GAN / VAE
2015 ─ ResNet / DQN / BatchNorm
2017 ─ Transformer / PPO / AlphaZero
2018 ─ BERT / GPT-1 ◀────────────────────────        The beginning of the scale era
2020 ─ GPT-3 / DDPM
2022 ─ InstructGPT (RLHF) / Latent Diffusion ◀        Generative AI goes mainstream
2023 ─ GPT-4, multimodal large models                  The era of general foundation models

Only two observations matter most from this timeline:

  • 2012 (AlexNet) and 2018 (BERT) are two watershed moments. Before 2012, the winning formula was "feature engineering"; after 2012, it's "representation that auto-learns features." After 2018, "scale + data + compute" replaced "task-specific models" as the new game. Full narrative is in Evolutionary History.
  • Every paradigm shift has not "eliminated" the previous era but demoted it to a sub-problem. Feature engineering becomes "a step in representation learning" (data preprocessing, prompt design); SVM/trees still live on the tabular data battlefield; rules return even in the form of "prompt constraints." On the map, there are no "eliminated papers," only "papers that changed battlefields."

4. Papers by Task ​

The map is organized by theme, but in practice you usually search for papers by task. Below re-indexes representative papers from the seven themes under "what problem do I need to solve":

Task I want to doRead these papers in orderWhat you'll gain
Image classificationAlexNet → ResNet → (follow with ViT-family papers for frontier)Complete evolution from convolutions to depth
Object detectionFaster R-CNN (arXiv:1506.01497) → YOLO (arXiv:1506.02640)Two-stage vs one-stage route debate
Image segmentationU-Net (arXiv:1505.04597) → (follow-up: Mask R-CNN, etc.)Encoder-decoder + skip-connection structural template
Machine translationSeq2Seq (arXiv:1409.3215) → Bahdanau Attention → TransformerHow "encode-decode" upgrades to "full attention"
Text/code generationGPT-2 → GPT-3 → InstructGPT (RLHF)Complete chain of generation + alignment
Image generationGAN → DDPM → Latent DiffusionEvolution from adversarial to diffusion
Recommendation systemsNeural Collaborative Filtering (arXiv:1708.05031) → (plus industry XGBoost/LightGBM for ranking)Neural-ization of collaborative filtering
Tabular data modelingRandom Forests → XGBoost → LightGBM/CatBoostFrom ensemble principles to engineering choices

Two reminders about the task table

First, a few papers here (Faster R-CNN, YOLO, U-Net, Seq2Seq, NCF) are not in the main table of Section 2, but they are all names you must know in this field—the task index fills the gap between "theme map" and "task practice." Second, a task does not equal a model: the optimal model for the same task varies by data shape (CNN-family for images, trees for tables)—refer to the data-type perspective in Deep Learning Foundations for model selection.

5. How to Use This Map ​

1. Deep-Read Core: Eight Must-Read Papers ​

The biggest mistake in paper reading is "spreading effort evenly." I recommend concentrating your limited time on eight "hub papers"—each is a rite of passage intersection on its technical line. After absorbing them, you'll be able to understand the abstracts of most other papers. In recommended order:

OrderPaperWhy Deep-Read It
1Attention Is All You NeedThe architectural source of all contemporary large models—reading it first is like getting the map key for the 2020s
2AlexNetThe detonation point of the deep learning revolution, short and intuitive, great for building the first feel of "how to read a paper"
3ResNetResidual connections influence every deep network—understanding it is essential to "why modern models can be hundreds of layers deep"
4BERTPretraining/fine-tuning paradigm + bidirectional attention—the turning point where NLP shifted to large models
5GPT-3Empirical foundation of scaling laws and in-context learning—reading it prepares you intellectually for ChatGPT
6DDPMThe origin of diffusion models—the common foundation of current image/video/audio generation technologies
7XGBoostOne of the most widely used papers in industry, and the best demonstration of "reading a paper → using the right tool"
8PPO or InstructGPT (pick one)Read PPO for deep RL, read InstructGPT to understand ChatGPT's "alignment"

Step-by-step deep reads of these eight are in Classic Paper Deep Dives; methodology on how to read (note-taking, what to do when you can't understand) is at Reading Discipline & FAQ.

2. Skim the Periphery: One Chain Per Theme ​

Not every paper deserves deep reading. For the remaining papers in the map of Section 2, use the "three-piece skim method": abstract (1 min) → key figure (2 min) → conclusion & limitations (2 min), producing three-line notes: "problem solved / core method in one sentence / self-stated limitations." This turns each of the seven themes into a "chain," and your map transforms from "seven islands" into "one network."

3. On-Demand Retrieval: Four-Step Method ​

When facing a specific problem (e.g., "I need to do object detection"), use the map in four steps:

① Locate the task → check Section 4 "Papers by Task" table, find the entry paper
② Trace the source → read the entry paper's References, find what it stands on (source tracing)
③ Trace downstream → use Google Scholar / Semantic Scholar to find "papers that cite it," find the latest methods (flow direction)
④ File the node → hang the new paper onto a theme in Section 2's map + a year on the timeline

After one round, your personal map will be larger than this article—this is exactly the correct use of a map: inherit first, then expand.

4. Coordinate with Other Resources on This Site ​

6. Further Reading ​

References ​

All materials below are real, publicly accessible resources for further self-study. All arXiv numbers can be directly accessed at https://arxiv.org/abs/<number>: