Skip to content

The Paper Map

At a glance With Attention Is All You Need as the root node, this page lays out four evolutionary branches — large language models, generative models, engineering applications, and alignment — into a searchable paper map covering 20+ key papers from 2014–2025, complete with paper-concept-case mappings and reading priorities.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

The Paper Map ​

In one sentence: this map takes the 2017 paper Attention Is All You Need as its root node and arranges the 20-plus key papers from 2014 to 2025 that shaped modern AI into a coordinate system you can consult at any time, organized along four evolutionary branches. Once you've worked through this map, you should be able to say, for every paper, which earlier work it stands on, which new branches it grew, and whether it deserves a close reading or just a skim.

The biggest obstacle in reading papers is not "failing to understand one" but "not knowing which one to read next". Tutorials teach in chapter order, but the real world of papers is organized by citation relationships: GPT-3 borrowed BERT's pretraining idea yet switched the goal from "understanding" to "generation"; DDPM and GAN competed on the same problem of "producing realistic images" yet took completely different roads; LoRA, the "put large models on a diet" solution, only appeared two years after GPT-3 shipped. If you read purely along the timeline, you get lost in the details of "who replaced whom"; if you read purely by topic, you miss the macro narrative of "why paradigms shift". The purpose of this map is to stack the two dimensions of topic and time on top of each other, so that every paper can answer three questions: which technical line does it belong to? At what stage of paradigm evolution does it sit? And which papers form a "must-read chain" with it?

How this page divides work with the rest of the site

The papers section of this site spans six pages, each with its own job: Start Here covers "why read papers at all"; Reading Paths covers "how to read given your goal"; this page (the paper map) covers "which papers in the whole field are worth reading and how they relate"; Classic Paper Deep-Dives digs into ten hinge papers one by one; Frontier Developments tracks the new trends of the 2020s; and Reading Discipline and FAQ answers "what if I can't understand, can't remember, or can't judge quality". In one line: the map fixes the coordinates, the deep-dives give depth, and the frontier points the way.

1. How to Read This Map ​

Before unfolding the map, three reading conventions:

  1. Numbering convention: except for GAN (NeurIPS 2014) and the Sora report (an OpenAI blog post, no arXiv), every paper in this article has an arXiv ID, and the full text is freely available at https://arxiv.org/abs/<id>.
  2. Year convention: a paper's year always means the year of first public release (formal conference publication or the arXiv preprint submission). For example, Latent Diffusion's arXiv preprint was submitted in December 2021 and published at CVPR 2022, so this article records it as 2021 (formally published 2022); BERT's arXiv submission was October 2018, with NAACL 2019 publication, so it is recorded as 2018. The labeling rule matters more than any specific number.
  3. Deep-read convention: the map only gives a "one-sentence contribution" and a "reading priority", answering "why is this worth remembering, and is it worth a deep dive". When you want to truly digest a paper, jump to Classic Paper Deep-Dives; to follow what came after it, use Frontier Developments; to decide what order to read in, go back to Reading Paths.

Use the three coordinates together: the branch tells you which technical line a paper belongs to, the year tells you where it sits in paradigm evolution, and the priority tells you how much time to invest. Here is the map at a glance:

                 Topic dimension (four evolutionary branches)
              ┌────────────────────────────────────────────────────┐
Time dim. ↓   │  ① Language models: understanding → scale → alignment → open source → reasoning │
              │  ② Generative models: adversarial → diffusion → control → video          │
              │  ③ Engineering: retrieval → fine-tuning → agents → knowledge graphs       │
              │  ④ Alignment: human feedback → AI feedback → preference optimization     │
              └────────────────────────────────────────────────────┘

2. The Root Node: Attention Is All You Need (2017) ​

Where do the four branches come from? They all grow from the root of the same tree — Attention Is All You Need, published by a Google team in 2017. It replaced both recurrence and convolution with a pure attention architecture (the Transformer); it started life as merely a new machine-translation model, yet within a few years it became the common foundation beneath large language models (LLMs), multimodal models, and diffusion-model backbones. It bears no fruit of its own — instead, it gave every successor the same ground to build on.

PaperAuthors/AffiliationYearVenue/arXivContribution in one sentence
Attention Is All You NeedVaswani et al. (Google Brain)2017arXiv:1706.03762Proposed the pure-attention Transformer, dropping recurrence and convolution — the shared foundation beneath BERT, GPT, and even diffusion-model backbones

From this root, the four branches each evolved:

                 ┌── Language models: BERT → GPT-3 → InstructGPT → ChatGPT → Llama → o1/R1
Attention       ├── Generative models: GAN → DDPM → LDM → ControlNet → Sora
Is All You Need ├── Engineering: RAG → LoRA → ReAct → QLoRA → GraphRAG
 (2017)         └── Alignment: InstructGPT → Constitutional AI → DPO

Why 2017 Is the Root

There were of course plenty of important works before the Transformer (Seq2Seq, the attention mechanism, CNNs/RNNs), but for today's hot AI concepts the Transformer is the "common genetic ancestor" — this site's brief history of AI tells the fuller story, and the Transformer concept page covers its mechanics. The map starts in 2017 because every branch that follows can be traced back here along citation lines.

3. The Four Evolutionary Branches ​

1. The language-model branch: from "understanding" to "scale" to "reasoning" (2018–2025) ​

This branch is the trunk of the 2020s AI boom. It answers three questions in sequence: how does a model learn to understand language (BERT) → does scale directly buy capability (GPT-3) → how do we make models more obedient and better at reasoning (InstructGPT / ChatGPT / Llama / o1 / R1). For the mechanics in depth, see Large Language Models.

PaperYearVenue/arXivContribution in one sentenceEvolutionary arrowPriority
BERT2018arXiv:1810.04805Bidirectional Transformer encoder + masked language model pretraining; swept 11 NLP benchmarks and established the "pretrain + fine-tune" paradigmGrew out of the Transformer (the encoder line)Must-read
GPT-32020arXiv:2005.14165175 billion parameters + in-context learning; demonstrated empirically that "scale alone brings capability jumps" and kicked off scaling laws← BERT's pretraining idea, swapped to an autoregressive generation objectiveMust-read
InstructGPT2022arXiv:2203.02155Used RLHF to teach GPT-3 to "follow instructions", aligning it with human intent← GPT-3 plus an "alignment" stepMust-read
ChatGPT2022No formal paper (OpenAI technical note)Packaged GPT-3.5 + RLHF into a conversational product, igniting generative AI's mainstream moment← InstructGPT turned into a productExtension
Llama2023arXiv:2302.13971Openly redistributable foundation models (the Llama family) that brought fine-tuning and local deployment to everyone← GPT-3's architecture recipe, taken open sourceOptional
o1 / DeepSeek-R12024–2025arXiv:2412.16735 (o1 system card) / arXiv:2501.12948 (R1)Test-time compute + reinforcement-learning-driven chain-of-thought — models learn to "think before answering"← InstructGPT's RL idea transferred to reasoningOptional

The Bidirectional vs. Unidirectional Fork

BERT and GPT set out from the same Transformer but took two opposite roads: BERT is a bidirectional encoder, good at understanding (classification, extraction, retrieval); GPT is an autoregressive decoder, good at generation. In the 2020s the contest was ultimately won by the generative road — ChatGPT proved that "treat every task as generation" is the more unified paradigm. This fork and convergence is the key thread for understanding AI history from 2018 to 2024; read it alongside a brief history of AI, and for the full product-level story see ChatGPT and Conversational AI.

2. The generative-model branch: from "adversarial" to "diffusion" to "video" (2014–2024) ​

This branch answers a question with broader cultural reach: can a machine, like a painter, produce an image — even a video — from a single sentence? How the answer changed is a history of paradigm shifts: GAN was the first to fool the eye through adversarial play; DDPM leapfrogged it with "add noise, then denoise", winning on stability and quality; LDM moved diffusion into latent space and turned image generation into a mass-market product; ControlNet handed creators precise control; and Sora extended the same recipe to minute-long video. For the mechanics in depth, see Diffusion Models and Generative AI.

PaperYearVenue/arXivContribution in one sentenceEvolutionary arrowPriority
GAN2014arXiv:1406.2661Generator vs. discriminator in a zero-sum game — the first model to make "photorealistic images from noise" a realityThe opening shot of generative AI (predates the Transformer)Optional
DDPM2020arXiv:2006.11239Progressively adds noise, then learns to denoise; stable training with no mode collapse — laid the foundation of diffusion-based generationTook over GAN's role as the main generative workhorseMust-read
LDM (Latent Diffusion)2021arXiv:2112.10752Diffusion in latent space + text-condition injection; efficient, controllable training — the foundation of Stable Diffusion← DDPM moved into latent spaceMust-read
ControlNet2023arXiv:2302.05543Adds "control handles" (pose, edges, depth) to diffusion models, so generation results can be directed precisely← Controllability enhancement for LDMOptional
Sora report2024Official blog (no arXiv)Diffusion Transformer + video data; claims video generation models are "world simulators"← LDM thinking extended to video; the preceding paper is DiT (arXiv:2212.09748)Extension

This Branch in One Sentence

GAN chased realism through adversarial play, DDPM chased stability through denoising, LDM chased efficiency through the latent space, ControlNet chased control through conditioning, and Sora chased world realism through video. This branch maps one-to-one onto Midjourney and Image Generation and Sora and Video Generation; anyone trying to put "text-to-image" into production will also want Prompt Engineering.

3. The engineering branch: putting large models to work (2020–2024) ​

The first two branches ask "how do models get stronger?"; this one asks "how do models get used?". RAG fixes "models can't remember the latest knowledge"; LoRA fixes "full fine-tuning is too expensive"; ReAct fixes "letting models actually do things"; QLoRA squeezes fine-tuning onto a single consumer GPU; and GraphRAG patches RAG's weak spot in global understanding with knowledge graphs. This branch is the closest to engineering practice; the hands-on companions are Build a RAG App from Scratch and Build an Agent from Scratch.

PaperYearVenue/arXivContribution in one sentenceEvolutionary arrowPriority
RAG2020arXiv:2005.11401Retrieves from an external knowledge base before generating; eases hallucinations and lets the model "know up-to-date facts"Same year as GPT-3; starts a separate engineering lineMust-read
LoRA2021arXiv:2106.09685Fine-tunes by training only low-rank delta matrices; parameter-efficient, and spawned the PEFT family← GPT-3's fine-tuning pain pointMust-read
ReAct2022arXiv:2210.03629Has the LLM alternate between "reasoning" and "acting" (calling tools); laid the foundation of the agent paradigm← The second tool-use route besides RAGMust-read
QLoRA2023arXiv:2305.143144-bit quantization + LoRA; fine-tunes a 65B model on a single GPU← LoRA compressed againOptional
GraphRAG2024arXiv:2404.16130Builds a knowledge graph from documents and then runs RAG over it; answering "global questions" no longer means stitching fragments← RAG + knowledge graphExtension

How the Engineering Branch Maps to This Site's Practice Pages

For RAG in production, see Retrieval-Augmented Generation (RAG), Vector Databases, and Perplexity and AI Search; for LoRA/QLoRA, see Fine-Tuning and PEFT and Fine-Tune Your Own LLM; for ReAct, see AI Agents, Manus and Agent Applications, and Build an Agent from Scratch. This is the branch where you can write code the moment you finish the paper.

4. The alignment branch: making models speak plainly and play by the rules (2022–2023) ​

The first three branches are about capability; this one is about values and obedience. InstructGPT established the RLHF trio (SFT + reward model + PPO); Constitutional AI tried replacing expensive human feedback with "AI reviewing AI"; and DPO achieves "imitate the preferred, reject the disliked" with one closed-form objective, reducing RLHF to a single supervised-learning run. For the conceptual walkthrough of this branch, see Alignment: RLHF and DPO; for the broader governance view, see AI Safety and Governance.

PaperYearVenue/arXivContribution in one sentenceEvolutionary arrowPriority
InstructGPT2022arXiv:2203.02155Human feedback + reinforcement learning (PPO as the engine) aligns a large model to user intent — the core technology behind ChatGPTOrigin of the alignment branch (shared with the language-model branch)Must-read
Constitutional AI2022arXiv:2212.08073Replaces part of the human feedback with "a constitution + AI self-critique", lowering alignment costs← InstructGPT's cost-reduction follow-upOptional
DPO2023arXiv:2305.18290Optimizes the policy directly from preference data, dropping the reward model and online sampling; radically simple training← A simplification of InstructGPTMust-read

A "Citation Line" That Is Easy to Misread

InstructGPT appears in both the language-model branch and the alignment branch — not because one paper is being hung on two hooks, but because one paper really did two things: in capability terms it turned GPT-3 into "an assistant that listens"; in method terms it established the RLHF paradigm. This is exactly where a map earns its keep: a paper can span multiple topics, and while reading it you should hang it on both lines at once.

4. The Paper–Concept–Case Mapping Table ​

A map is only as good as its use. The table below maps every paper in this article to this site's concept pages (learn the principles) and case-study pages (see the products), with priorities marked. Whenever a paper's name comes up at work, you can follow this table to the matching deep content on this site.

PaperConcept pageCase studyPriority
Attention Is All You NeedTransformer and the Attention MechanismChatGPT and Conversational AIMust-read
BERTLarge Language ModelsPerplexity and AI SearchMust-read
GPT-3Large Language ModelsChatGPT and Conversational AIMust-read
InstructGPTAlignment: RLHF and DPOChatGPT and Conversational AIMust-read
ChatGPTLarge Language ModelsChatGPT and Conversational AIExtension
LlamaFine-Tuning and PEFTDeepSeek-R1 and Reasoning ModelsOptional
o1 / DeepSeek-R1Large Language ModelsDeepSeek-R1 and Reasoning ModelsOptional
GANDiffusion Models and Generative AIMidjourney and Image GenerationOptional
DDPMDiffusion Models and Generative AIMidjourney and Image GenerationMust-read
LDMDiffusion Models and Generative AIMidjourney and Image GenerationMust-read
ControlNetDiffusion Models and Generative AIMidjourney and Image GenerationOptional
Sora reportMultimodal ModelsSora and Video GenerationExtension
RAGRetrieval-Augmented GenerationPerplexity and AI SearchMust-read
LoRAFine-Tuning and PEFTGitHub Copilot and Code IntelligenceMust-read
ReActAI AgentsManus and Agent ApplicationsMust-read
QLoRAFine-Tuning and PEFTGitHub Copilot and Code IntelligenceOptional
GraphRAGKnowledge Graphs and Knowledge InjectionPerplexity and AI SearchExtension
Constitutional AIAlignment: RLHF and DPOChatGPT and Conversational AIOptional
DPOAlignment: RLHF and DPODeepSeek-R1 and Reasoning ModelsMust-read

How to Read the Priority Labels

Must-read: read it thoroughly and you can understand the abstracts of most papers in that branch — these are the "hinge papers"; Optional: read it after the must-reads and it completes a full chain; Extension: mostly technical reports or product notes — skim them alongside Frontier Developments. If you only have time for 8 papers, follow the order of the ten must-reads in section 5.

5. How to Use This Map ​

1. The deep-read core: ten must-reads ​

The worst mistake in paper reading is spreading your effort evenly. Concentrate your limited time on the ten must-read "hinge papers" — each one is the unavoidable crossroads of some technical line. In recommended order:

OrderPaperWhy it deserves a close read
1Attention Is All You NeedThe architectural source of every modern large model; reading it first hands you the key to the 2020s map
2BERTThe pretrain/fine-tune paradigm + bidirectional attention — understand "how models learn language"
3GPT-3The foundational evidence for scaling laws and in-context learning; reading it prepares you conceptually for ChatGPT
4InstructGPTThe RLHF trio; understand "why models obey", and step into the alignment branch
5DDPMThe origin of diffusion models — the shared root of today's image/video/audio generation
6LDMThe foundation of Stable Diffusion; understand "why image generation runs on consumer GPUs"
7RAGOne of the most widely deployed engineering techniques; the first lesson in hallucination control and bolting on external knowledge
8LoRAThe representative of parameter-efficient fine-tuning; nearly the whole open-source model ecosystem uses it
9ReActFounded the agent paradigm; "reasoning + acting" moves models from answering questions to doing things
10DPOThe second paradigm of alignment; radically simple training, and part of the alignment family behind models like R1

For paper-by-paper deep readings of these ten, see Classic Paper Deep-Dives; for reading methodology (how to take notes, what to do when you don't understand), see Reading Discipline and FAQ.

2. Skimming the periphery: one chain per branch ​

Not every paper deserves a close read. For Optional/Extension papers, use the three-piece skim: abstract (1 minute) → key figure (2 minutes) → conclusions and limitations (2 minutes), producing a three-line note: "what problem it solves / the core method in one sentence / the authors' own stated limitations". Read each of the four branches as a "chain" this way, and your map turns from "four isolated islands" into "one connected network".

3. Retrieval by task: the four-step method ​

When you hit a concrete problem (say, "I need to hook a knowledge base up to a large model"), use the map in four steps:

① Locate the branch  → decide whether the problem belongs to language models / generation / engineering / alignment, and find the entry paper
② Trace upstream     → read the entry paper's References to find which earlier work it stands on (where it came from)
③ Trace downstream   → use Google Scholar / Semantic Scholar to find "papers citing it" (where it flows)
④ Record the node    → attach the new paper to some branch + some year in this document, growing your own map

After one round, your personal map is bigger than this article — which is exactly how a map is meant to be used: inherit first, then extend.

4. Pairing this map with the site's other resources ​

Further Reading ​

References ​

All of the resources below are real, publicly available materials for further self-study. Every arXiv ID can be accessed directly at https://arxiv.org/abs/<id>: