Theme
Paper Map
One-sentence summary: This page is a "map of LLM papers" — laid out vertically by timeline and sliced horizontally by topic, so you can see at a glance which papers are origins, which are milestones, and which form the same technical tributary. With a coordinate system in hand, you won't get lost in the sea of 2,000 arXiv papers.
1. Why You Need a Map
The biggest cost of reading papers isn't "reading" — it's "deciding which paper to read." A map solves three problems for you:
- Know where the origins are: For example, LoRA's roots trace back to the "low-rank approximation" idea and Adapter (2019). Understanding origins lets you judge whether a new paper is innovative or rehashing old ideas.
- Know how papers inherit from each other: GPT-3 depends on GPT-2 and the Scaling Laws belief; InstructGPT depends on RLHF; DPO is a simplification of RLHF. The map draws out this "citation chain."
- Know your own position: If you're in applications, focus on the "applications and systems" tributary; if you're in alignment, focus on the "alignment era."
The narrative version of this map (without paper details) is on Evolution Timeline. Reading both pages together works best.
2. Map Overview: Six Eras
| Era | Rough Time | Topic Keywords | One-Line Summary | Corresponding Section |
|---|---|---|---|---|
| Pre-Transformer | 1950s–2016 | Statistics, representations, recurrent | From n-gram to neural representations, attention appears as a "plugin" | Evolution Timeline |
| Architecture Revolution | 2017–2018 | Self-attention, pretraining | Transformer emerges; GPT and BERT establish two pretraining routes | Transformer Architecture |
| Scale Era | 2019–2021 | Scaling laws, few-shot | Parameters and data grow exponentially; "bigger is better" becomes dogma | Scaling Laws |
| Alignment Era | 2022–2023 | RLHF, instruction following | From "can generate" to "obeys instructions"; ChatGPT goes mainstream | Alignment |
| Applications & Systems | 2021–2024 | Fine-tuning, retrieval, acceleration | Making LLMs "usable": LoRA, RAG, FlashAttention | RAG |
| Frontier | 2024–2025 | Inference scaling, native multimodal | Inference-time scaling, long context, MoE scaling, open-source catching up to closed-source | Frontier Trends |
3. Map by Era
3.1 Pre-Transformer Era (1950s–2016)
| Year | Paper / Work | One-Line Contribution | Significance |
|---|---|---|---|
| 1948 | Shannon's A Mathematical Theory of Communication | Foundations of information theory; "predicting the next symbol" becomes the theoretical root of language modeling | Origin |
| 1980s–90s | n-gram statistical language models | Estimate probability using word-sequence frequencies; Markov assumption | Origin |
| 2003 | Bengio et al. A Neural Probabilistic Language Model | Neural networks learn word representations and conditional probability; birth of neural language modeling | Milestone |
| 2013 | Mikolov et al. Distributed Representations of Words and Phrases | Word2Vec: word embedding training technique; explosion of representation learning | Milestone |
| 2014 | Sutskever et al. Sequence to Sequence Learning | Encoder–decoder framework; end-to-end sequence modeling | Milestone |
| 2014 | Bahdanau et al. Neural Machine Translation by Jointly Learning to Align | First to use attention (alignment) for translation | Milestone |
| 2015–2016 | RNN/LSTM language models and machine translation systems | Recurrent networks dominate, but can't parallelize and struggle with long-range memory | Background |
One-line takeaway for this era
The core tension of this era was "sequences can't parallelize + memory window is too short." Transformer solved both problems at once — so the moment it arrived, the era turned.
Era summary: This era's legacy is two ideas — "representation learning" (Word2Vec) and "sequence modeling" (Seq2Seq, attention). Attention was just an auxiliary plugin in translation systems at the time. The lesson was clear: sequences are hard to parallelize, and memory windows are too short. Transformer's later success came precisely from tearing down both walls simultaneously.
3.2 Architecture Revolution (2017–2018)
| Year | Paper | One-Line Contribution | Significance |
|---|---|---|---|
| 2017 | Vaswani et al. Attention Is All You Need | Transformer: self-attention + multi-head + positional encoding; parallelizable, long-range | Origin (everything starts here) |
| 2018 | Radford et al. Improving Language Understanding by Generative Pre-Training (GPT-1) | Generative pretraining + fine-tuning; opens the decoder route | Milestone |
| 2018 | Peters et al. Deep Contextualized Word Representations (ELMo) | Contextual word embeddings; bidirectional LSTM | Transition |
| 2018 | Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers | Masked language model + bidirectional encoder; dominance in understanding tasks | Milestone |
| 2019 | Dai et al. Transformer-XL | Segment-level recurrence + relative positional encoding; early long-context attempt | Tributary |
The divergence point of two routes
Decoder route (GPT): unidirectional, autoregressive, excels at generation, supports few-shot; Encoder route (BERT): bidirectional, masked, excels at understanding, requires fine-tuning. After 2023, the decoder route basically won — but BERT's legacy (bidirectional representations, embeddings, retrieval) still lives on. See GPT Series and BERT and the Encoder Family.
Era summary: Transformer's contribution wasn't just "a new architecture" — it simultaneously lit up two paradigms: "pretraining + fine-tuning" and "self-attention." The GPT–BERT split (unidirectional decoder vs. bidirectional encoder) determined the five-year route battle, and the two pretraining objectives (autoregressive vs. masked) remain core concepts in language modeling.
3.3 Scale Era (2019–2021)
| Year | Paper | One-Line Contribution | Significance |
|---|---|---|---|
| 2019 | Radford et al. Language Models are Unsupervised Multitask Learners (GPT-2) | 1.5B params + WebText; proves zero-shot transfer | Milestone |
| 2019 | Liu et al. RoBERTa | Squeezes more out of BERT with more data and longer training | Engineering |
| 2019 | Raffel et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) | Unified "text-to-text" paradigm; encoder–decoder representative | Tributary |
| 2020 | Kaplan et al. Scaling Laws for Neural Language Models | Loss decreases as a power law of params/data/compute; the "theoretical license" for scaling | Milestone |
| 2020 | Brown et al. Language Models are Few-Shot Learners (GPT-3) | 175B + few-shot; empirical manifesto of scaling laws | Milestone |
| 2020 | Lewis et al. Retrieval-Augmented Generation (RAG) | Retrieval-augmented generation; external knowledge injection | Milestone (applications branch) |
| 2020 | Lepikhin et al. GShard | Conditionally activated sparse expert models; MoE foundation | Origin (MoE) |
| 2020 | Dosovitskiy et al. An Image is Worth 16x16 Words (ViT) | Transformer unifies vision | Tributary |
| 2021 | Fedus et al. Switch Transformers | Simplified MoE routing; trillion-parameter training becomes feasible | Milestone (MoE) |
| 2021 | Chen et al. Evaluating Large Language Models Trained on Code (Codex) | Code generation + GitHub Copilot's foundation | Milestone (applications) |
| 2021 | Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP) | Image-text contrastive learning; early multimodal paradigm | Origin (multimodal) |
| 2021 | Hu et al. LoRA: Low-Rank Adaptation of Large Language Models | Parameter-efficient fine-tuning via low-rank updates | Milestone (applications) |
Era summary: Scaling laws turned the intuition "bigger is better" into "a quantitatively predictable power law." GPT-3 empirically demonstrated few-shot capability with 175B parameters. Meanwhile, three tributaries were laid out: MoE (GShard/Switch) laid the foundation for sparsity, ViT brought Transformer to vision, and Codex proved code is excellent training corpus. The era's mantra was "scale up" — more parameters, more data, more compute.
3.4 Alignment Era (2022–2023)
| Year | Paper / Work | One-Line Contribution | Significance |
|---|---|---|---|
| 2022 | Ouyang et al. Training Language Models to Follow Instructions with Human Feedback (InstructGPT) | SFT → RM → PPO, three-step RLHF to make models "obedient" | Milestone |
| 2022 | Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | Chain-of-thought prompting, elicits reasoning ability | Milestone |
| 2022 | Hoffmann et al. Training Compute-Optimal Large Language Models (Chinchilla) | Compute-optimal ratio (~1:20), corrects "bigger is always better" | Milestone |
| 2022 | Wang et al. Self-Consistency Improves Chain of Thought Reasoning | Majority vote aggregates multiple CoT traces | Tributary |
| 2022 | Bai et al. Constitutional AI | AI feedback as a substitute for human feedback in alignment | Tributary |
| 2022 | Wang et al. Self-Instruct | Let the model self-generate instruction data | Tributary |
| 2022.11 | ChatGPT | InstructGPT technology + conversational product; mainstream breakthrough | Milestone (product) |
| 2023 | Ouyang team + OpenAI GPT-4 Technical Report | Multimodal, exam-level; industry capability ceiling | Milestone |
| 2023 | Touvron et al. Llama: Open and Efficient Foundation Language Models | Open-source model returns, community ecosystem ignition point | Milestone |
| 2023 | Rafailov et al. Direct Preference Optimization (DPO) | Preference learning without a reward model, simplifies RLHF | Milestone |
| 2023 | Dettmers et al. QLoRA | 4-bit quantization + LoRA; fine-tune 65B on a single GPU | Milestone (engineering) |
| 2023 | Touvron et al. Llama 2 | Open-source + commercial license + conversational model weights | Milestone |
| 2023 | Jiang et al. Mistral 7B | Engineering victory of small models competing with large ones | Tributary |
| 2023 | Jiang et al. Mixtral of Experts | Open-source sparse MoE (8×7B) | Milestone (MoE) |
Era summary: InstructGPT proved that "can generate" and "can obey" are two different things — RLHF became the default alignment option. ChatGPT productized the technology and went mainstream, while Llama open-sourced weights and ignited the community. The era's keyword shifted from "capability" to "controllability," with the cost of introducing an "alignment tax" (some benchmark scores dropped slightly after alignment) — until later people realized that alignment and capability could mutually reinforce each other.
3.5 Applications and Systems (2021–2024)
| Year | Paper | One-Line Contribution | Significance |
|---|---|---|---|
| 2022 | Dao et al. FlashAttention | I/O-aware exact attention, memory from O(n²) to O(n) | Milestone (systems) |
| 2022 | Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models | Reasoning + acting interleaved; Agent paradigm foundation | Milestone (Agent) |
| 2023 | Schick et al. Toolformer | Let models learn to call tools themselves | Tributary (Agent) |
| 2023 | Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) | Paged KV Cache, order-of-magnitude throughput boost | Milestone (systems) |
| 2023 | Peng et al. LongLoRA | Efficient context extension (e.g., 4k→32k) | Tributary (long context) |
| 2023 | Peng et al. YaRN | Practical RoPE length extrapolation method | Tributary (long context) |
| 2023 | Dao et al. FlashAttention-2 | Parallelization and better tiling, another ~2× speedup | Systems |
| 2023 | Hong et al. MetaGPT | Multi-agent collaborative software company | Tributary (Agent) |
Era summary: Once models and alignment were in place, "affordable to use, fast to run" became the main thread: LoRA democratized fine-tuning, FlashAttention made long context feasible, PagedAttention (vLLM) boosted inference throughput by an order of magnitude, and ReAct established the Agent paradigm. This era reminds us: model papers determine the ceiling, system papers determine the floor — most engineering bottlenecks are in systems, not models.
3.6 Frontier (2024–2025)
| Year | Paper / Work | One-Line Contribution | Significance |
|---|---|---|---|
| 2023.12 | Gu & Dao Mamba: Linear-Time Sequence Modeling with Selective State Spaces | State space model, linear complexity, hardware-friendly | Milestone (new architecture direction) |
| 2024.02 | Gemini 1.5 Technical Report | Million-level context window goes live | Milestone (long context) |
| 2024.03 | OpenAI Sora Technical Report | Video generation as a world simulator | Tributary (multimodal) |
| 2024.03 | AI21 Jamba | Mamba + Transformer hybrid architecture, open-source | Tributary |
| 2024.05 | OpenAI GPT-4o | Native multimodal + real-time interaction | Milestone (product) |
| 2024 | Touvron et al. / Meta Llama 3 | 8B/70B/405B, 15T tokens, open-source catches up to closed-source | Milestone (open-source) |
| 2024.09 | OpenAI o1 | Inference-time scaling: chain of thought + reinforcement learning | Milestone (reasoning paradigm) |
| 2024.12 | DeepSeek DeepSeek-V3 | 671B total / 37B activated, ultra-low training cost | Milestone (MoE) |
| 2025.01 | DeepSeek DeepSeek-R1 | Pure RL training of reasoning model, open-source alternative to o1-level capability | Milestone (reasoning + open-source) |
Era summary: The keyword for 2024–2025 is "fourth-dimensional scaling": inference-time scaling (o1) trades inference compute for capability, extending Scaling Laws from "training compute" to "inference compute"; MoE drives costs down via sparse activation; and open-source models use "smaller, more data, more efficient training" to approach or even match closed-source. The frontier row of this map is being rewritten on a monthly basis.
The map is a "snapshot as of mid-2025"
The frontier row changes daily. The map's value is in its structure (six eras, four thematic tributaries), not the timeliness of the latest entry. See Frontier Trends and the "following arXiv" section of Reading Discipline & FAQ for ways to stay current.
4. Topic Slices: Four Major Tributaries
If you look by "topic" instead of by "time," papers form four tributaries. When deep-diving, it's recommended to follow tributaries rather than scanning by era:
| Topic Tributary | Key Papers (chronological) | One-Line Narrative | Site Coverage |
|---|---|---|---|
| Architecture | Attention → GPT-1 → BERT → GPT-2 → FlashAttention → Mamba | From "attention is all you need" to "attention + state-space hybrid" | Transformer, Deep-dives |
| Scale | Scaling Laws → GPT-3 → Chinchilla → MoE series | From "more params = better" to "compute optimal" to "sparse activation" | Scaling Laws, MoE |
| Alignment | RLHF (2017) → InstructGPT → Constitutional AI → DPO → o1-style RL | From "make models obedient" to "make models reason" | Alignment |
| Applications & Systems | RAG → LoRA → ReAct → vLLM → long context | Make LLMs "affordable, usable, and fast" | RAG, Multimodal |
5. Surveys and Further Resources
Every tributary on the map has a corresponding survey paper — the fastest way to "stand on giants' shoulders and scan the map":
- A Survey of Large Language Models (Zhao et al., 2023) — LLM panorama survey: five major blocks covering architecture, pretraining, adaptation, evaluation, and safety. Great for absorbing the whole-page map at once.
- A Survey of Transformers (Tay et al., 2022) — The full Transformer family: attention variants, architecture variants, applications. An encyclopedia of the architecture tributary.
- A Survey on Evaluation of Large Language Models (2023) — Evaluation survey: task classification, benchmark inventory, evaluation methods. The entry point for the evaluation tributary.
Additionally, there are recent survey papers on specialized topics like retrieval-augmented generation, Agents, and reasoning. Search arXiv for "survey + keyword" — prioritize versions updated within the last six months with high citation counts.
How to use surveys
Surveys are meant for "reading first," not "deep-reading": spend an afternoon going through a survey, paint the map in your mind, then go back to specific papers as needed. Note that surveys have inherent staleness, and the survey authors may have biases — always return to the original papers for final judgment.
6. Four Big Takeaways From the Map
After reading the whole map, here are four cross-era takeaways worth remembering:
The decoder route ultimately won. GPT and BERT were neck-and-neck in 2018, but by 2023 the decoder dominated — even understanding task benchmarks (like MT-Bench, IFEval) are ruled by decoder models. The bidirectional representation legacy didn't disappear — it entered retrieval encoders, embeddings, and evaluation (see the Encoder family).
Capability comes from pretraining scale; alignment organizes that capability. InstructGPT's 1.3B beat GPT-3's 175B not because alignment created capability, but because it organized the 175B's existing capability better. Without scale, there's nothing to align — which explains why alignment research always anchors on the largest models.
System papers determine whether capability can be deployed. Papers like FlashAttention and PagedAttention aren't as glamorous as "new models," but long context, low-cost inference, and single-GPU fine-tuning all rest on them. Don't only stare at the "model" tributaries when looking at the map.
Open-source closes the research loop. After Llama, every breakthrough almost has an open-source follow-up. The community went from "only read papers" to "read papers + run models + reproduce." DeepSeek-R1 even turned "inference-time scaling" — a closed-source direction — into an open-source benchmark.
7. How to Maintain Your Own Map
A map isn't static. We recommend doing one "map operation" every time you read a new paper:
text
□ Locate: Which era and which tributary does this paper belong to? (check this site's map)
□ Classify: Origin / Milestone / Tributary / Survey?
□ Relationships: What does it inherit from? What does it inspire? (check citations and references)
□ Update: If it's a milestone, add it to your own "frontier" row
□ Archive: Hang the card note under the corresponding tributary (see [card note method](/papers/faq))Personal map = this site's map + your own notes. In three months, you'll have your own "history of LLM technological evolution."
8. From Map to Paths
- Go wide then deep: Spend 10 minutes on the Section 2 overview first, identify which tributary you're in.
- Pick milestones: For each tributary, pick 2–3 papers labeled "milestone" and deep-dive them (checklist in Reading Paths).
- Follow the vine: After reading a milestone, continue expanding via the Related Work and citation chains in its text — the map is alive, the more you read, the denser it gets.
- Cross-reference with deep-dive pages: Core Paper Deep Dives has already deep-dived the most important nodes on the map for you — use it as a "coordinate calibrator."
Further Reading
- Reading Paths — After selecting papers from the map, get the reading order by goal
- Core Paper Deep Dives — In-depth readings of the map's most important nodes
- Frontier Trends — Latest additions to the map's "frontier" row, organized by trend
- Evolution Timeline — The narrative version of the map
- GPT Series: From GPT-1 to GPT-4o — The main-thread case from architecture revolution to alignment era
References
Key papers referenced on this page (real arXiv links, grouped by topic):
- Bengio et al. A Neural Probabilistic Language Model (2003) — JMLR official version
- Mikolov et al. Distributed Representations of Words and Phrases (2013) — Word2Vec
- Bahdanau et al. Neural Machine Translation by Jointly Learning to Align and Translate (2014) — Attention mechanism origin
- Sutskever et al. Sequence to Sequence Learning with Neural Networks (2014) — Encoder–decoder framework
- Peters et al. Deep Contextualized Word Representations (2018) — ELMo
- Liu et al. RoBERTa (2019) — Stronger BERT
- Raffel et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2019) — T5
- Dai et al. Transformer-XL (2019) — Early long-context attempt
- Fedus et al. Switch Transformers (2021) — Trillion-parameter MoE
- Chen et al. Evaluating Large Language Models Trained on Code (2021) — Codex
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (2021) — CLIP
- Touvron et al. Llama (2023) — Llama 1
- Jiang et al. Mistral 7B (2023) — Mistral 7B
- Kwon et al. PagedAttention / vLLM (2023) — Paged KV Cache
- Peng et al. LongLoRA (2023) — Efficient context extension
- Touvron et al. Llama 2 (2023) — Llama 2
- Team Gemini. Gemini 1.5: Unlocking multimodal understanding across millions of tokens (2024) — Million-token context
- Gu & Dao. Mamba (2023) — State space model
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024) — MoE at massive scale
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025) — Open-source sample of inference-time scaling