Appearance
The Paper Map
In one sentence: this map takes the 2017 paper Attention Is All You Need as its root node and arranges the 20-plus key papers from 2014 to 2025 that shaped modern AI into a coordinate system you can consult at any time, organized along four evolutionary branches. Once you've worked through this map, you should be able to say, for every paper, which earlier work it stands on, which new branches it grew, and whether it deserves a close reading or just a skim.
The biggest obstacle in reading papers is not "failing to understand one" but "not knowing which one to read next". Tutorials teach in chapter order, but the real world of papers is organized by citation relationships: GPT-3 borrowed BERT's pretraining idea yet switched the goal from "understanding" to "generation"; DDPM and GAN competed on the same problem of "producing realistic images" yet took completely different roads; LoRA, the "put large models on a diet" solution, only appeared two years after GPT-3 shipped. If you read purely along the timeline, you get lost in the details of "who replaced whom"; if you read purely by topic, you miss the macro narrative of "why paradigms shift". The purpose of this map is to stack the two dimensions of topic and time on top of each other, so that every paper can answer three questions: which technical line does it belong to? At what stage of paradigm evolution does it sit? And which papers form a "must-read chain" with it?
How this page divides work with the rest of the site
The papers section of this site spans six pages, each with its own job: Start Here covers "why read papers at all"; Reading Paths covers "how to read given your goal"; this page (the paper map) covers "which papers in the whole field are worth reading and how they relate"; Classic Paper Deep-Dives digs into ten hinge papers one by one; Frontier Developments tracks the new trends of the 2020s; and Reading Discipline and FAQ answers "what if I can't understand, can't remember, or can't judge quality". In one line: the map fixes the coordinates, the deep-dives give depth, and the frontier points the way.
1. How to Read This Map
Before unfolding the map, three reading conventions:
- Numbering convention: except for GAN (NeurIPS 2014) and the Sora report (an OpenAI blog post, no arXiv), every paper in this article has an arXiv ID, and the full text is freely available at
https://arxiv.org/abs/<id>. - Year convention: a paper's year always means the year of first public release (formal conference publication or the arXiv preprint submission). For example, Latent Diffusion's arXiv preprint was submitted in December 2021 and published at CVPR 2022, so this article records it as 2021 (formally published 2022); BERT's arXiv submission was October 2018, with NAACL 2019 publication, so it is recorded as 2018. The labeling rule matters more than any specific number.
- Deep-read convention: the map only gives a "one-sentence contribution" and a "reading priority", answering "why is this worth remembering, and is it worth a deep dive". When you want to truly digest a paper, jump to Classic Paper Deep-Dives; to follow what came after it, use Frontier Developments; to decide what order to read in, go back to Reading Paths.
Use the three coordinates together: the branch tells you which technical line a paper belongs to, the year tells you where it sits in paradigm evolution, and the priority tells you how much time to invest. Here is the map at a glance:
Topic dimension (four evolutionary branches)
┌────────────────────────────────────────────────────┐
Time dim. ↓ │ ① Language models: understanding → scale → alignment → open source → reasoning │
│ ② Generative models: adversarial → diffusion → control → video │
│ ③ Engineering: retrieval → fine-tuning → agents → knowledge graphs │
│ ④ Alignment: human feedback → AI feedback → preference optimization │
└────────────────────────────────────────────────────┘2. The Root Node: Attention Is All You Need (2017)
Where do the four branches come from? They all grow from the root of the same tree — Attention Is All You Need, published by a Google team in 2017. It replaced both recurrence and convolution with a pure attention architecture (the Transformer); it started life as merely a new machine-translation model, yet within a few years it became the common foundation beneath large language models (LLMs), multimodal models, and diffusion-model backbones. It bears no fruit of its own — instead, it gave every successor the same ground to build on.
| Paper | Authors/Affiliation | Year | Venue/arXiv | Contribution in one sentence |
|---|---|---|---|---|
| Attention Is All You Need | Vaswani et al. (Google Brain) | 2017 | arXiv:1706.03762 | Proposed the pure-attention Transformer, dropping recurrence and convolution — the shared foundation beneath BERT, GPT, and even diffusion-model backbones |
From this root, the four branches each evolved:
┌── Language models: BERT → GPT-3 → InstructGPT → ChatGPT → Llama → o1/R1
Attention ├── Generative models: GAN → DDPM → LDM → ControlNet → Sora
Is All You Need ├── Engineering: RAG → LoRA → ReAct → QLoRA → GraphRAG
(2017) └── Alignment: InstructGPT → Constitutional AI → DPOWhy 2017 Is the Root
There were of course plenty of important works before the Transformer (Seq2Seq, the attention mechanism, CNNs/RNNs), but for today's hot AI concepts the Transformer is the "common genetic ancestor" — this site's brief history of AI tells the fuller story, and the Transformer concept page covers its mechanics. The map starts in 2017 because every branch that follows can be traced back here along citation lines.
3. The Four Evolutionary Branches
1. The language-model branch: from "understanding" to "scale" to "reasoning" (2018–2025)
This branch is the trunk of the 2020s AI boom. It answers three questions in sequence: how does a model learn to understand language (BERT) → does scale directly buy capability (GPT-3) → how do we make models more obedient and better at reasoning (InstructGPT / ChatGPT / Llama / o1 / R1). For the mechanics in depth, see Large Language Models.
| Paper | Year | Venue/arXiv | Contribution in one sentence | Evolutionary arrow | Priority |
|---|---|---|---|---|---|
| BERT | 2018 | arXiv:1810.04805 | Bidirectional Transformer encoder + masked language model pretraining; swept 11 NLP benchmarks and established the "pretrain + fine-tune" paradigm | Grew out of the Transformer (the encoder line) | Must-read |
| GPT-3 | 2020 | arXiv:2005.14165 | 175 billion parameters + in-context learning; demonstrated empirically that "scale alone brings capability jumps" and kicked off scaling laws | ← BERT's pretraining idea, swapped to an autoregressive generation objective | Must-read |
| InstructGPT | 2022 | arXiv:2203.02155 | Used RLHF to teach GPT-3 to "follow instructions", aligning it with human intent | ← GPT-3 plus an "alignment" step | Must-read |
| ChatGPT | 2022 | No formal paper (OpenAI technical note) | Packaged GPT-3.5 + RLHF into a conversational product, igniting generative AI's mainstream moment | ← InstructGPT turned into a product | Extension |
| Llama | 2023 | arXiv:2302.13971 | Openly redistributable foundation models (the Llama family) that brought fine-tuning and local deployment to everyone | ← GPT-3's architecture recipe, taken open source | Optional |
| o1 / DeepSeek-R1 | 2024–2025 | arXiv:2412.16735 (o1 system card) / arXiv:2501.12948 (R1) | Test-time compute + reinforcement-learning-driven chain-of-thought — models learn to "think before answering" | ← InstructGPT's RL idea transferred to reasoning | Optional |
The Bidirectional vs. Unidirectional Fork
BERT and GPT set out from the same Transformer but took two opposite roads: BERT is a bidirectional encoder, good at understanding (classification, extraction, retrieval); GPT is an autoregressive decoder, good at generation. In the 2020s the contest was ultimately won by the generative road — ChatGPT proved that "treat every task as generation" is the more unified paradigm. This fork and convergence is the key thread for understanding AI history from 2018 to 2024; read it alongside a brief history of AI, and for the full product-level story see ChatGPT and Conversational AI.
2. The generative-model branch: from "adversarial" to "diffusion" to "video" (2014–2024)
This branch answers a question with broader cultural reach: can a machine, like a painter, produce an image — even a video — from a single sentence? How the answer changed is a history of paradigm shifts: GAN was the first to fool the eye through adversarial play; DDPM leapfrogged it with "add noise, then denoise", winning on stability and quality; LDM moved diffusion into latent space and turned image generation into a mass-market product; ControlNet handed creators precise control; and Sora extended the same recipe to minute-long video. For the mechanics in depth, see Diffusion Models and Generative AI.
| Paper | Year | Venue/arXiv | Contribution in one sentence | Evolutionary arrow | Priority |
|---|---|---|---|---|---|
| GAN | 2014 | arXiv:1406.2661 | Generator vs. discriminator in a zero-sum game — the first model to make "photorealistic images from noise" a reality | The opening shot of generative AI (predates the Transformer) | Optional |
| DDPM | 2020 | arXiv:2006.11239 | Progressively adds noise, then learns to denoise; stable training with no mode collapse — laid the foundation of diffusion-based generation | Took over GAN's role as the main generative workhorse | Must-read |
| LDM (Latent Diffusion) | 2021 | arXiv:2112.10752 | Diffusion in latent space + text-condition injection; efficient, controllable training — the foundation of Stable Diffusion | ← DDPM moved into latent space | Must-read |
| ControlNet | 2023 | arXiv:2302.05543 | Adds "control handles" (pose, edges, depth) to diffusion models, so generation results can be directed precisely | ← Controllability enhancement for LDM | Optional |
| Sora report | 2024 | Official blog (no arXiv) | Diffusion Transformer + video data; claims video generation models are "world simulators" | ← LDM thinking extended to video; the preceding paper is DiT (arXiv:2212.09748) | Extension |
This Branch in One Sentence
GAN chased realism through adversarial play, DDPM chased stability through denoising, LDM chased efficiency through the latent space, ControlNet chased control through conditioning, and Sora chased world realism through video. This branch maps one-to-one onto Midjourney and Image Generation and Sora and Video Generation; anyone trying to put "text-to-image" into production will also want Prompt Engineering.
3. The engineering branch: putting large models to work (2020–2024)
The first two branches ask "how do models get stronger?"; this one asks "how do models get used?". RAG fixes "models can't remember the latest knowledge"; LoRA fixes "full fine-tuning is too expensive"; ReAct fixes "letting models actually do things"; QLoRA squeezes fine-tuning onto a single consumer GPU; and GraphRAG patches RAG's weak spot in global understanding with knowledge graphs. This branch is the closest to engineering practice; the hands-on companions are Build a RAG App from Scratch and Build an Agent from Scratch.
| Paper | Year | Venue/arXiv | Contribution in one sentence | Evolutionary arrow | Priority |
|---|---|---|---|---|---|
| RAG | 2020 | arXiv:2005.11401 | Retrieves from an external knowledge base before generating; eases hallucinations and lets the model "know up-to-date facts" | Same year as GPT-3; starts a separate engineering line | Must-read |
| LoRA | 2021 | arXiv:2106.09685 | Fine-tunes by training only low-rank delta matrices; parameter-efficient, and spawned the PEFT family | ← GPT-3's fine-tuning pain point | Must-read |
| ReAct | 2022 | arXiv:2210.03629 | Has the LLM alternate between "reasoning" and "acting" (calling tools); laid the foundation of the agent paradigm | ← The second tool-use route besides RAG | Must-read |
| QLoRA | 2023 | arXiv:2305.14314 | 4-bit quantization + LoRA; fine-tunes a 65B model on a single GPU | ← LoRA compressed again | Optional |
| GraphRAG | 2024 | arXiv:2404.16130 | Builds a knowledge graph from documents and then runs RAG over it; answering "global questions" no longer means stitching fragments | ← RAG + knowledge graph | Extension |
How the Engineering Branch Maps to This Site's Practice Pages
For RAG in production, see Retrieval-Augmented Generation (RAG), Vector Databases, and Perplexity and AI Search; for LoRA/QLoRA, see Fine-Tuning and PEFT and Fine-Tune Your Own LLM; for ReAct, see AI Agents, Manus and Agent Applications, and Build an Agent from Scratch. This is the branch where you can write code the moment you finish the paper.
4. The alignment branch: making models speak plainly and play by the rules (2022–2023)
The first three branches are about capability; this one is about values and obedience. InstructGPT established the RLHF trio (SFT + reward model + PPO); Constitutional AI tried replacing expensive human feedback with "AI reviewing AI"; and DPO achieves "imitate the preferred, reject the disliked" with one closed-form objective, reducing RLHF to a single supervised-learning run. For the conceptual walkthrough of this branch, see Alignment: RLHF and DPO; for the broader governance view, see AI Safety and Governance.
| Paper | Year | Venue/arXiv | Contribution in one sentence | Evolutionary arrow | Priority |
|---|---|---|---|---|---|
| InstructGPT | 2022 | arXiv:2203.02155 | Human feedback + reinforcement learning (PPO as the engine) aligns a large model to user intent — the core technology behind ChatGPT | Origin of the alignment branch (shared with the language-model branch) | Must-read |
| Constitutional AI | 2022 | arXiv:2212.08073 | Replaces part of the human feedback with "a constitution + AI self-critique", lowering alignment costs | ← InstructGPT's cost-reduction follow-up | Optional |
| DPO | 2023 | arXiv:2305.18290 | Optimizes the policy directly from preference data, dropping the reward model and online sampling; radically simple training | ← A simplification of InstructGPT | Must-read |
A "Citation Line" That Is Easy to Misread
InstructGPT appears in both the language-model branch and the alignment branch — not because one paper is being hung on two hooks, but because one paper really did two things: in capability terms it turned GPT-3 into "an assistant that listens"; in method terms it established the RLHF paradigm. This is exactly where a map earns its keep: a paper can span multiple topics, and while reading it you should hang it on both lines at once.
4. The Paper–Concept–Case Mapping Table
A map is only as good as its use. The table below maps every paper in this article to this site's concept pages (learn the principles) and case-study pages (see the products), with priorities marked. Whenever a paper's name comes up at work, you can follow this table to the matching deep content on this site.
How to Read the Priority Labels
Must-read: read it thoroughly and you can understand the abstracts of most papers in that branch — these are the "hinge papers"; Optional: read it after the must-reads and it completes a full chain; Extension: mostly technical reports or product notes — skim them alongside Frontier Developments. If you only have time for 8 papers, follow the order of the ten must-reads in section 5.
5. How to Use This Map
1. The deep-read core: ten must-reads
The worst mistake in paper reading is spreading your effort evenly. Concentrate your limited time on the ten must-read "hinge papers" — each one is the unavoidable crossroads of some technical line. In recommended order:
| Order | Paper | Why it deserves a close read |
|---|---|---|
| 1 | Attention Is All You Need | The architectural source of every modern large model; reading it first hands you the key to the 2020s map |
| 2 | BERT | The pretrain/fine-tune paradigm + bidirectional attention — understand "how models learn language" |
| 3 | GPT-3 | The foundational evidence for scaling laws and in-context learning; reading it prepares you conceptually for ChatGPT |
| 4 | InstructGPT | The RLHF trio; understand "why models obey", and step into the alignment branch |
| 5 | DDPM | The origin of diffusion models — the shared root of today's image/video/audio generation |
| 6 | LDM | The foundation of Stable Diffusion; understand "why image generation runs on consumer GPUs" |
| 7 | RAG | One of the most widely deployed engineering techniques; the first lesson in hallucination control and bolting on external knowledge |
| 8 | LoRA | The representative of parameter-efficient fine-tuning; nearly the whole open-source model ecosystem uses it |
| 9 | ReAct | Founded the agent paradigm; "reasoning + acting" moves models from answering questions to doing things |
| 10 | DPO | The second paradigm of alignment; radically simple training, and part of the alignment family behind models like R1 |
For paper-by-paper deep readings of these ten, see Classic Paper Deep-Dives; for reading methodology (how to take notes, what to do when you don't understand), see Reading Discipline and FAQ.
2. Skimming the periphery: one chain per branch
Not every paper deserves a close read. For Optional/Extension papers, use the three-piece skim: abstract (1 minute) → key figure (2 minutes) → conclusions and limitations (2 minutes), producing a three-line note: "what problem it solves / the core method in one sentence / the authors' own stated limitations". Read each of the four branches as a "chain" this way, and your map turns from "four isolated islands" into "one connected network".
3. Retrieval by task: the four-step method
When you hit a concrete problem (say, "I need to hook a knowledge base up to a large model"), use the map in four steps:
① Locate the branch → decide whether the problem belongs to language models / generation / engineering / alignment, and find the entry paper
② Trace upstream → read the entry paper's References to find which earlier work it stands on (where it came from)
③ Trace downstream → use Google Scholar / Semantic Scholar to find "papers citing it" (where it flows)
④ Record the node → attach the new paper to some branch + some year in this document, growing your own mapAfter one round, your personal map is bigger than this article — which is exactly how a map is meant to be used: inherit first, then extend.
4. Pairing this map with the site's other resources
- When a concept is fuzzy, check the Glossary;
- For structured knowledge of a concept, read Concept Boundaries and Overall Architecture Anatomy;
- To see the papers described here as real products, read ChatGPT and Conversational AI, Midjourney and Image Generation, and Perplexity and AI Search;
- To get hands-on, follow the routes in the practice guide, or go straight to Common Pitfalls and Antipatterns;
- To see what job requirements these papers correspond to, check the JD knowledge-point breakdown.
Further Reading
- Classic Paper Deep-Dives — paper-by-paper close readings of the ten must-reads on this map (fact sheet, motivation, method, results, impact, pitfalls)
- Frontier Developments — the right end of the map's timeline: new papers and new trends of 2024–2025
- A Brief History of AI — the complete historical narrative behind the paradigm-shift timeline
- Start Here and Reading Paths — orientation and route choice before you start reading
- Reading Discipline and FAQ — come here when you can't understand, can't remember, or can't judge a paper's quality
- Glossary — look up concepts on the fly while reading
References
All of the resources below are real, publicly available materials for further self-study. Every arXiv ID can be accessed directly at https://arxiv.org/abs/<id>:
- Vaswani et al. Attention Is All You Need. NeurIPS 2017 — the Transformer, root node of this map
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2019 — BERT
- Brown et al. Language Models are Few-Shot Learners. NeurIPS 2020 — GPT-3
- Ouyang et al. Training language models to follow instructions with human feedback. NeurIPS 2022 — InstructGPT / RLHF
- Touvron et al. LLaMA: Open and Efficient Foundation Language Models. 2023 — Llama 1
- OpenAI. OpenAI o1 System Card. 2024 — the o1 system card
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025 — DeepSeek-R1
- Goodfellow et al. Generative Adversarial Nets. NeurIPS 2014 — GAN
- Ho, Jain, Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS 2020 — DDPM
- Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022 — LDM / the foundation of Stable Diffusion
- Zhang et al. Adding Conditional Control to Text-to-Image Diffusion Models. ICLR 2024 — ControlNet
- Peebles & Xie. Scalable Diffusion Models with Transformers. ICCV 2023 — DiT, the architectural predecessor of Sora
- OpenAI. Video generation models as world simulators. 2024 — the Sora technical report
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 — RAG
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022 — LoRA
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 — ReAct
- Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023 — QLoRA
- Edge et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. 2024 — GraphRAG
- Bai et al. Constitutional AI: Harmlessness from AI Feedback. 2022 — Constitutional AI
- Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023 — DPO
- The arXiv preprint library — first-release home of the vast majority of the papers in this article
- Papers with Code — papers + code + benchmarks in one place; the authoritative entry point for checking SOTA and reproductions