Skip to content

Paper Map

At a glance A map of key LLM papers over decades, arranged by timeline and topic: pre-Transformer era, architecture revolution, scale era, alignment era, applications and systems, and the frontier. Each paper gets one row with year, one-line contribution, and significance.

Paper Map ​

One-sentence summary: This page is a "map of LLM papers" — laid out vertically by timeline and sliced horizontally by topic, so you can see at a glance which papers are origins, which are milestones, and which form the same technical tributary. With a coordinate system in hand, you won't get lost in the sea of 2,000 arXiv papers.

1. Why You Need a Map ​

The biggest cost of reading papers isn't "reading" — it's "deciding which paper to read." A map solves three problems for you:

  1. Know where the origins are: For example, LoRA's roots trace back to the "low-rank approximation" idea and Adapter (2019). Understanding origins lets you judge whether a new paper is innovative or rehashing old ideas.
  2. Know how papers inherit from each other: GPT-3 depends on GPT-2 and the Scaling Laws belief; InstructGPT depends on RLHF; DPO is a simplification of RLHF. The map draws out this "citation chain."
  3. Know your own position: If you're in applications, focus on the "applications and systems" tributary; if you're in alignment, focus on the "alignment era."

The narrative version of this map (without paper details) is on Evolution Timeline. Reading both pages together works best.

2. Map Overview: Six Eras ​

EraRough TimeTopic KeywordsOne-Line SummaryCorresponding Section
Pre-Transformer1950s–2016Statistics, representations, recurrentFrom n-gram to neural representations, attention appears as a "plugin"Evolution Timeline
Architecture Revolution2017–2018Self-attention, pretrainingTransformer emerges; GPT and BERT establish two pretraining routesTransformer Architecture
Scale Era2019–2021Scaling laws, few-shotParameters and data grow exponentially; "bigger is better" becomes dogmaScaling Laws
Alignment Era2022–2023RLHF, instruction followingFrom "can generate" to "obeys instructions"; ChatGPT goes mainstreamAlignment
Applications & Systems2021–2024Fine-tuning, retrieval, accelerationMaking LLMs "usable": LoRA, RAG, FlashAttentionRAG
Frontier2024–2025Inference scaling, native multimodalInference-time scaling, long context, MoE scaling, open-source catching up to closed-sourceFrontier Trends

3. Map by Era ​

3.1 Pre-Transformer Era (1950s–2016) ​

YearPaper / WorkOne-Line ContributionSignificance
1948Shannon's A Mathematical Theory of CommunicationFoundations of information theory; "predicting the next symbol" becomes the theoretical root of language modelingOrigin
1980s–90sn-gram statistical language modelsEstimate probability using word-sequence frequencies; Markov assumptionOrigin
2003Bengio et al. A Neural Probabilistic Language ModelNeural networks learn word representations and conditional probability; birth of neural language modelingMilestone
2013Mikolov et al. Distributed Representations of Words and PhrasesWord2Vec: word embedding training technique; explosion of representation learningMilestone
2014Sutskever et al. Sequence to Sequence LearningEncoder–decoder framework; end-to-end sequence modelingMilestone
2014Bahdanau et al. Neural Machine Translation by Jointly Learning to AlignFirst to use attention (alignment) for translationMilestone
2015–2016RNN/LSTM language models and machine translation systemsRecurrent networks dominate, but can't parallelize and struggle with long-range memoryBackground

One-line takeaway for this era

The core tension of this era was "sequences can't parallelize + memory window is too short." Transformer solved both problems at once — so the moment it arrived, the era turned.

Era summary: This era's legacy is two ideas — "representation learning" (Word2Vec) and "sequence modeling" (Seq2Seq, attention). Attention was just an auxiliary plugin in translation systems at the time. The lesson was clear: sequences are hard to parallelize, and memory windows are too short. Transformer's later success came precisely from tearing down both walls simultaneously.

3.2 Architecture Revolution (2017–2018) ​

YearPaperOne-Line ContributionSignificance
2017Vaswani et al. Attention Is All You NeedTransformer: self-attention + multi-head + positional encoding; parallelizable, long-rangeOrigin (everything starts here)
2018Radford et al. Improving Language Understanding by Generative Pre-Training (GPT-1)Generative pretraining + fine-tuning; opens the decoder routeMilestone
2018Peters et al. Deep Contextualized Word Representations (ELMo)Contextual word embeddings; bidirectional LSTMTransition
2018Devlin et al. BERT: Pre-training of Deep Bidirectional TransformersMasked language model + bidirectional encoder; dominance in understanding tasksMilestone
2019Dai et al. Transformer-XLSegment-level recurrence + relative positional encoding; early long-context attemptTributary

The divergence point of two routes

Decoder route (GPT): unidirectional, autoregressive, excels at generation, supports few-shot; Encoder route (BERT): bidirectional, masked, excels at understanding, requires fine-tuning. After 2023, the decoder route basically won — but BERT's legacy (bidirectional representations, embeddings, retrieval) still lives on. See GPT Series and BERT and the Encoder Family.

Era summary: Transformer's contribution wasn't just "a new architecture" — it simultaneously lit up two paradigms: "pretraining + fine-tuning" and "self-attention." The GPT–BERT split (unidirectional decoder vs. bidirectional encoder) determined the five-year route battle, and the two pretraining objectives (autoregressive vs. masked) remain core concepts in language modeling.

3.3 Scale Era (2019–2021) ​

YearPaperOne-Line ContributionSignificance
2019Radford et al. Language Models are Unsupervised Multitask Learners (GPT-2)1.5B params + WebText; proves zero-shot transferMilestone
2019Liu et al. RoBERTaSqueezes more out of BERT with more data and longer trainingEngineering
2019Raffel et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)Unified "text-to-text" paradigm; encoder–decoder representativeTributary
2020Kaplan et al. Scaling Laws for Neural Language ModelsLoss decreases as a power law of params/data/compute; the "theoretical license" for scalingMilestone
2020Brown et al. Language Models are Few-Shot Learners (GPT-3)175B + few-shot; empirical manifesto of scaling lawsMilestone
2020Lewis et al. Retrieval-Augmented Generation (RAG)Retrieval-augmented generation; external knowledge injectionMilestone (applications branch)
2020Lepikhin et al. GShardConditionally activated sparse expert models; MoE foundationOrigin (MoE)
2020Dosovitskiy et al. An Image is Worth 16x16 Words (ViT)Transformer unifies visionTributary
2021Fedus et al. Switch TransformersSimplified MoE routing; trillion-parameter training becomes feasibleMilestone (MoE)
2021Chen et al. Evaluating Large Language Models Trained on Code (Codex)Code generation + GitHub Copilot's foundationMilestone (applications)
2021Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP)Image-text contrastive learning; early multimodal paradigmOrigin (multimodal)
2021Hu et al. LoRA: Low-Rank Adaptation of Large Language ModelsParameter-efficient fine-tuning via low-rank updatesMilestone (applications)

Era summary: Scaling laws turned the intuition "bigger is better" into "a quantitatively predictable power law." GPT-3 empirically demonstrated few-shot capability with 175B parameters. Meanwhile, three tributaries were laid out: MoE (GShard/Switch) laid the foundation for sparsity, ViT brought Transformer to vision, and Codex proved code is excellent training corpus. The era's mantra was "scale up" — more parameters, more data, more compute.

3.4 Alignment Era (2022–2023) ​

YearPaper / WorkOne-Line ContributionSignificance
2022Ouyang et al. Training Language Models to Follow Instructions with Human Feedback (InstructGPT)SFT → RM → PPO, three-step RLHF to make models "obedient"Milestone
2022Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsChain-of-thought prompting, elicits reasoning abilityMilestone
2022Hoffmann et al. Training Compute-Optimal Large Language Models (Chinchilla)Compute-optimal ratio (~1:20), corrects "bigger is always better"Milestone
2022Wang et al. Self-Consistency Improves Chain of Thought ReasoningMajority vote aggregates multiple CoT tracesTributary
2022Bai et al. Constitutional AIAI feedback as a substitute for human feedback in alignmentTributary
2022Wang et al. Self-InstructLet the model self-generate instruction dataTributary
2022.11ChatGPTInstructGPT technology + conversational product; mainstream breakthroughMilestone (product)
2023Ouyang team + OpenAI GPT-4 Technical ReportMultimodal, exam-level; industry capability ceilingMilestone
2023Touvron et al. Llama: Open and Efficient Foundation Language ModelsOpen-source model returns, community ecosystem ignition pointMilestone
2023Rafailov et al. Direct Preference Optimization (DPO)Preference learning without a reward model, simplifies RLHFMilestone
2023Dettmers et al. QLoRA4-bit quantization + LoRA; fine-tune 65B on a single GPUMilestone (engineering)
2023Touvron et al. Llama 2Open-source + commercial license + conversational model weightsMilestone
2023Jiang et al. Mistral 7BEngineering victory of small models competing with large onesTributary
2023Jiang et al. Mixtral of ExpertsOpen-source sparse MoE (8×7B)Milestone (MoE)

Era summary: InstructGPT proved that "can generate" and "can obey" are two different things — RLHF became the default alignment option. ChatGPT productized the technology and went mainstream, while Llama open-sourced weights and ignited the community. The era's keyword shifted from "capability" to "controllability," with the cost of introducing an "alignment tax" (some benchmark scores dropped slightly after alignment) — until later people realized that alignment and capability could mutually reinforce each other.

3.5 Applications and Systems (2021–2024) ​

YearPaperOne-Line ContributionSignificance
2022Dao et al. FlashAttentionI/O-aware exact attention, memory from O(n²) to O(n)Milestone (systems)
2022Yao et al. ReAct: Synergizing Reasoning and Acting in Language ModelsReasoning + acting interleaved; Agent paradigm foundationMilestone (Agent)
2023Schick et al. ToolformerLet models learn to call tools themselvesTributary (Agent)
2023Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM)Paged KV Cache, order-of-magnitude throughput boostMilestone (systems)
2023Peng et al. LongLoRAEfficient context extension (e.g., 4k→32k)Tributary (long context)
2023Peng et al. YaRNPractical RoPE length extrapolation methodTributary (long context)
2023Dao et al. FlashAttention-2Parallelization and better tiling, another ~2× speedupSystems
2023Hong et al. MetaGPTMulti-agent collaborative software companyTributary (Agent)

Era summary: Once models and alignment were in place, "affordable to use, fast to run" became the main thread: LoRA democratized fine-tuning, FlashAttention made long context feasible, PagedAttention (vLLM) boosted inference throughput by an order of magnitude, and ReAct established the Agent paradigm. This era reminds us: model papers determine the ceiling, system papers determine the floor — most engineering bottlenecks are in systems, not models.

3.6 Frontier (2024–2025) ​

YearPaper / WorkOne-Line ContributionSignificance
2023.12Gu & Dao Mamba: Linear-Time Sequence Modeling with Selective State SpacesState space model, linear complexity, hardware-friendlyMilestone (new architecture direction)
2024.02Gemini 1.5 Technical ReportMillion-level context window goes liveMilestone (long context)
2024.03OpenAI Sora Technical ReportVideo generation as a world simulatorTributary (multimodal)
2024.03AI21 JambaMamba + Transformer hybrid architecture, open-sourceTributary
2024.05OpenAI GPT-4oNative multimodal + real-time interactionMilestone (product)
2024Touvron et al. / Meta Llama 38B/70B/405B, 15T tokens, open-source catches up to closed-sourceMilestone (open-source)
2024.09OpenAI o1Inference-time scaling: chain of thought + reinforcement learningMilestone (reasoning paradigm)
2024.12DeepSeek DeepSeek-V3671B total / 37B activated, ultra-low training costMilestone (MoE)
2025.01DeepSeek DeepSeek-R1Pure RL training of reasoning model, open-source alternative to o1-level capabilityMilestone (reasoning + open-source)

Era summary: The keyword for 2024–2025 is "fourth-dimensional scaling": inference-time scaling (o1) trades inference compute for capability, extending Scaling Laws from "training compute" to "inference compute"; MoE drives costs down via sparse activation; and open-source models use "smaller, more data, more efficient training" to approach or even match closed-source. The frontier row of this map is being rewritten on a monthly basis.

The map is a "snapshot as of mid-2025"

The frontier row changes daily. The map's value is in its structure (six eras, four thematic tributaries), not the timeliness of the latest entry. See Frontier Trends and the "following arXiv" section of Reading Discipline & FAQ for ways to stay current.

4. Topic Slices: Four Major Tributaries ​

If you look by "topic" instead of by "time," papers form four tributaries. When deep-diving, it's recommended to follow tributaries rather than scanning by era:

Topic TributaryKey Papers (chronological)One-Line NarrativeSite Coverage
ArchitectureAttention → GPT-1 → BERT → GPT-2 → FlashAttention → MambaFrom "attention is all you need" to "attention + state-space hybrid"Transformer, Deep-dives
ScaleScaling Laws → GPT-3 → Chinchilla → MoE seriesFrom "more params = better" to "compute optimal" to "sparse activation"Scaling Laws, MoE
AlignmentRLHF (2017) → InstructGPT → Constitutional AI → DPO → o1-style RLFrom "make models obedient" to "make models reason"Alignment
Applications & SystemsRAG → LoRA → ReAct → vLLM → long contextMake LLMs "affordable, usable, and fast"RAG, Multimodal

5. Surveys and Further Resources ​

Every tributary on the map has a corresponding survey paper — the fastest way to "stand on giants' shoulders and scan the map":

Additionally, there are recent survey papers on specialized topics like retrieval-augmented generation, Agents, and reasoning. Search arXiv for "survey + keyword" — prioritize versions updated within the last six months with high citation counts.

How to use surveys

Surveys are meant for "reading first," not "deep-reading": spend an afternoon going through a survey, paint the map in your mind, then go back to specific papers as needed. Note that surveys have inherent staleness, and the survey authors may have biases — always return to the original papers for final judgment.

6. Four Big Takeaways From the Map ​

After reading the whole map, here are four cross-era takeaways worth remembering:

  1. The decoder route ultimately won. GPT and BERT were neck-and-neck in 2018, but by 2023 the decoder dominated — even understanding task benchmarks (like MT-Bench, IFEval) are ruled by decoder models. The bidirectional representation legacy didn't disappear — it entered retrieval encoders, embeddings, and evaluation (see the Encoder family).

  2. Capability comes from pretraining scale; alignment organizes that capability. InstructGPT's 1.3B beat GPT-3's 175B not because alignment created capability, but because it organized the 175B's existing capability better. Without scale, there's nothing to align — which explains why alignment research always anchors on the largest models.

  3. System papers determine whether capability can be deployed. Papers like FlashAttention and PagedAttention aren't as glamorous as "new models," but long context, low-cost inference, and single-GPU fine-tuning all rest on them. Don't only stare at the "model" tributaries when looking at the map.

  4. Open-source closes the research loop. After Llama, every breakthrough almost has an open-source follow-up. The community went from "only read papers" to "read papers + run models + reproduce." DeepSeek-R1 even turned "inference-time scaling" — a closed-source direction — into an open-source benchmark.

7. How to Maintain Your Own Map ​

A map isn't static. We recommend doing one "map operation" every time you read a new paper:

text
□ Locate: Which era and which tributary does this paper belong to? (check this site's map)
□ Classify: Origin / Milestone / Tributary / Survey?
□ Relationships: What does it inherit from? What does it inspire? (check citations and references)
□ Update: If it's a milestone, add it to your own "frontier" row
□ Archive: Hang the card note under the corresponding tributary (see [card note method](/papers/faq))

Personal map = this site's map + your own notes. In three months, you'll have your own "history of LLM technological evolution."

8. From Map to Paths ​

  1. Go wide then deep: Spend 10 minutes on the Section 2 overview first, identify which tributary you're in.
  2. Pick milestones: For each tributary, pick 2–3 papers labeled "milestone" and deep-dive them (checklist in Reading Paths).
  3. Follow the vine: After reading a milestone, continue expanding via the Related Work and citation chains in its text — the map is alive, the more you read, the denser it gets.
  4. Cross-reference with deep-dive pages: Core Paper Deep Dives has already deep-dived the most important nodes on the map for you — use it as a "coordinate calibrator."

Further Reading ​

References ​

Key papers referenced on this page (real arXiv links, grouped by topic):