Theme
Reading Paths
One-sentence summary: This page is a decision table for "what to read, in what order, and to what depth." It organizes scattered arXiv papers into three actionable paths, so whether you have two hours or plan to do research, you can find your starting point here.
1. Two Principles to Keep in Mind
1. Go Wide First, Then Deep
Read "broadly covering" milestone papers first to build your coordinate system (see Paper Map), then pick one or two for deep dives (see Core Paper Deep Dives). Jumping straight into a 2025 paper without context will leave you both confused and wasting the opportunity to build a global view.
Why does this order work? Because new papers assume you're "someone who knows the context": they assume you've read Transformer, know RLHF, and recognize MoE. Without a horizontal coordinate system, every sentence in a new paper references something you don't know. With a coordinate system, a new paper's incremental information is only a thin layer. The cost of reading a paper is "context," and context comes from horizontal accumulation.
2. Define "How Deep" for Each Paper
For the same paper, beginners read abstract + conclusion, engineers read method + experiments, and researchers read derivations + reproduction. Setting a depth target for each paper is more efficient than powering through the full text. The full depth tiering (the three-pass reading method) is in Reading Discipline & FAQ. Here's a quick heuristic:
| Your Goal | Recommended Depth | Corresponding Pass |
|---|---|---|
| Understand trends, be able to chat | Abstract + intro + conclusion | First pass |
| Judge whether to apply to your work | Method + experiments + ablation + limitations | Second pass |
| Propose improvements, do research | Full text + derivations + reproduction | Third pass |
2. Essential Reading List (~12 Papers)
The table below is the site's "essential core" covering the full pipeline from architecture to systems to alignment. Note: By spec count, this merges to ~12 papers; listed individually here it's 13 rows (RLHF and DPO are grouped as one entry). The deep-dive versions are on Core Paper Deep Dives.
| # | Paper (Year) | Why It's Essential (One Sentence) | Sections to Focus On | What You Can Answer After Reading |
|---|---|---|---|---|
| 1 | Attention Is All You Need (2017) | The origin of everything Transformer: self-attention + multi-head + positional encoding | Abstract, §3 Model Architecture, §5 Experiments & Results | Why can self-attention replace RNN? |
| 2 | GPT-1 (2018) | First "generative pretraining + fine-tuning," opens the decoder route | Abstract, §3 Framework, Table 1 Results | Why does pretrain+fine-tune save labeled data? |
| 3 | GPT-2 (2019) | Proves "pretrained models can zero-shot transfer," introduces scale belief | Abstract, §2 Method, §4 Results | How does zero-shot happen "for free"? |
| 4 | BERT (2019) | Bidirectional encoder + masked language modeling, dominant paradigm for understanding tasks | Abstract, §3 Pre-training Objectives, §4 Ablations | Why is bidirectional more important for understanding? |
| 5 | GPT-3 (2020) | 175B parameters + few-shot, an empirical manifesto for scaling | Abstract, §2 Method, §3.9 Limitations, §3 Results | How does few-shot performance scale with model size? |
| 6 | Scaling Laws (2020) | Loss decreases as a power law of params/data/compute, guides "where to spend money" | Abstract, §3 Power Laws, §4 Transfer, Figures 1–4 | Should I add parameters or data? |
| 7 | Chinchilla (2022) | Compute-optimal ratio: params to tokens ≈ 1:20, corrects "bigger is always better" | Abstract, §4 Optimal Ratio, Table 3 Comparison | Why did a 70B model beat a 280B model? |
| 8 | InstructGPT (2022) | RLHF three-step method, the technical mother of ChatGPT | Abstract, §3 Method (SFT/RM/PPO), §4 Human Evaluation | How do models go from "can talk" to "obeys instructions"? |
| 9 | RLHF (2017) & DPO (2023) | A pair: preference learning evolves from "training a reward model" to "using preferences directly" | RLHF: §2 Method; DPO: §3 Derivation, Table 1 | Why can DPO skip the reward model? |
| 10 | LoRA (2021) | Parameter-efficient fine-tuning via low-rank updates, 10,000× reduction in fine-tuning cost | Abstract, §3.1 Low-Rank Parameterization, §4 Experiment Table | Why does the low-rank assumption hold? |
| 11 | FlashAttention (2022) | I/O-aware attention optimization, memory from O(n²) to linear, prerequisite for long context | Abstract, §3 Algorithm, §5 Speedup Table | Is the attention bottleneck compute or memory bandwidth? |
| 12 | Chain-of-Thought (2022) | The key to LLM reasoning ability, a watershed for prompting | Abstract, §2 Method, GSM8K Table, §4 Ablations | Why doesn't CoT work for small models? |
| 13 | RAG (2020) | The RAG paradigm origin, classic solution for hallucination mitigation and knowledge injection | Abstract, §3 Models (RAG-sequence/token), §4 Experiments | Why is an external "retrieval" better than "memorization"? |
How to use this table
You don't need to read them in numbered order. Beginners: start with 1, 5, 8, 10; Engineers: focus on 10, 11, 13; Researchers: read 6, 7, 9 thoroughly. Internal order within each path is in Section 6's dependency diagram.
3. Three Paths: Specific Checklists and Reading Order
Path A: 2-Hour Quick Start
Goal: Build the skeleton of "how Transformer works + why LLMs are powerful + how to make them obedient," so you can talk about it and not get lost in blogs.
text
Step 1 (40 min): Attention Is All You Need → Read only abstract, intro, conclusion + view The Illustrated Transformer visuals
Step 2 (30 min): GPT-1 → GPT-2 abstract comparison, understand "decoder + pretraining" route
Step 3 (30 min): GPT-3 abstract + limitations section, understand "scale + few-shot"
Step 4 (20 min): InstructGPT abstract + Figure 1, understand RLHF's three steps
Deliverable: Can say one sentence each for "self-attention, causal mask, pretraining, few-shot, RLHF"Requirement for this path: You don't need to understand formulas, but you must be able to answer three sentences per paper: "problem / method / effect." If a concept blocks you (e.g., causal mask), check Transformer Architecture or the glossary — don't struggle through it.
Path B: Engineering Implementation
Goal: Be able to judge "whether a paper's method can be applied to my work," and actually fine-tune or deploy it.
text
Prerequisite: Complete Path A first
Step 1: LoRA deep-dive method + experiments → cross-reference with Practice page [Fine-Tuning Practice: Full LoRA Workflow](/practice/fine-tuning-practice)
Step 2: RAG paper deep-dive §3–§4 → cross-reference with [case-studies RAG](/case-studies/rag) and [RAG in Practice](/practice/rag-in-practice)
Step 3: FlashAttention read abstract + algorithm → understand long context and inference speedup (paired with [Inference Fundamentals](/concepts/inference-fundamentals))
Step 4: Chinchilla read abstract + optimal ratio → understand "should I add data or parameters first?"
Step 5: CoT read method + GSM8K results → prompting engineering basis (paired with [Prompting](/concepts/prompting))
Deliverable: Can select "fine-tune vs. RAG vs. prompting" for a project and justify with paper evidenceRequirement for this path: For each paper, read "ablation + hyperparameters + limitations" and write a "risk checklist for transfer to my scenario." After finishing each paper, we recommend also checking the corresponding practice page in the practice guide — papers give principles, practice pages give the gotchas.
Path C: Research / Frontier Tracking
Goal: Stand at the 2025 frontier, read the latest arXiv, propose and validate improvements.
text
Prerequisite: Complete Paths A + B
Step 1: Read Scaling Laws full text (including appendix) and Chinchilla §4 derivation
Step 2: RLHF original paper (Christiano 2017) → InstructGPT → DPO, progressive reading to understand the evolution of preference learning
Step 3: Pick your own direction (alignment / systems / applications) and trace citation chains backwards from [Frontier Trends](/papers/frontier)
Step 4: Reproduce one paper (recommend a minimal implementation of LoRA or FlashAttention), build a "read + write" loop
Deliverable: Can write a one-page related work and propose a verifiable improvement hypothesisRequirement for this path: Reading papers isn't the end goal — writing is. Every third-pass deep-dive must produce: a mechanism explanation, a list of untested experiments, and an improvement hypothesis. Even if you don't validate the hypothesis, write it down — in three months you'll be surprised at how far you've come.
Comparison of the Three Paths
| Dimension | Path A: Intro | Path B: Engineering | Path C: Research |
|---|---|---|---|
| Target Audience | Beginners, product / non-ML roles | ML engineers, application devs | Grad students, researchers |
| Total Time | 2 hours | 2–3 weeks (1 hour/day) | Ongoing |
| Deep-Dive Count | 4 papers at abstract level | 6–8 papers deep | 15+ papers + continuous tracking |
| Key Actions | View visuals, grasp concepts | Reproduce experiments, write assessments | Derivations, reproduction, propose hypotheses |
| Deliverables | Three-sentence notes | Transfer risk checklist | Related work + improvement hypotheses |
| Validation Criteria | Can answer "why the GPT series succeeded" | Can deploy a feature based on a paper method | Can write related work, can defend your work |
| Corresponding Pages | Core Paper Deep Dives | Deep-dives + Frontier Trends + practice modules | Paper Map + Frontier Trends |
4. Paper Selection Quick-Reference by Goal Scenario
Not sure where to start? Look up by your current goal:
| Your Scenario | Read These First | Paired Pages |
|---|---|---|
| Want to understand Transformer and attention | 1, 11 | Transformer Architecture |
| Want to understand "why bigger models are stronger" | 6, 7, 5 | Scaling Laws |
| Want to do fine-tuning (SFT / LoRA) | 10, 2 | Fine-Tuning: SFT and PEFT, Fine-Tuning Practice |
| Want to understand alignment and RLHF / DPO | 8, 9 | Alignment: RLHF and DPO |
| Want to do RAG or reduce hallucination | 13, 12 | RAG Case Studies, RAG in Practice |
| Want to optimize inference / deployment performance | 11, 3 | Inference Fundamentals, Deployment Practice |
| Want to do evaluation / benchmarks | 6, 7 | Evaluation and Benchmarks, Evals in Practice |
| Preparing for ML interviews | 1, 5, 8, 10, 11, 13 deep | Interview Question Bank |
5. Generic Template: "Which Sections to Read in Each Paper"
Paper structures vary, but most follow an IMRaD structure (Introduction–Method–Results–Discussion). A generic trade-off and selection table:
| Section | Beginner Reader | Engineer Reader | Researcher Reader |
|---|---|---|---|
| Abstract + figure summaries | Required | Required | Required |
| Introduction | Required | Required | Required (focus on "gap" argument) |
| Related Work | Can skip | Skim | Required (follow citation chains) |
| Method | Figure only | Deep-dive + formulas | Deep-dive + derivations |
| Experiments | Main table only | Deep-dive ablations + hyperparams | Deep-dive + compare against reproduction |
| Limitations / Discussion | Required | Required (judge transferability) | Required (find improvement points) |
Specific selections for the essential reading list are already in the "reading approach" column of the Section 2 table. The full three-pass reading flow is in Reading Discipline & FAQ.
Why "Limitations" is required for everyone
Many readers skip limitations — and miss the most valuable information. The limitations section typically contains: under what conditions the method fails, what experiments the authors didn't do, and hints for future work — these are respectively the boundaries for engineering transfer, sources of improvement ideas, and clues for trend judgment.
6. Dependencies: Recommended Order
There's a clear inheritance chain between papers — reading in dependency order pays dividends:
text
Attention Is All You Need ─┬─→ GPT-1 → GPT-2 → GPT-3 ─┬─→ InstructGPT → DPO
└─→ BERT ──────────────────┘
↓
Scaling Laws ←─ Chinchilla ←─ InstructGPT (to understand "the cost of alignment")
Attention ──→ FlashAttention (to understand "how to speed up the same architecture")
GPT-3 ──→ RAG (to understand "parametric knowledge vs. retrieved knowledge")
GPT-3 ──→ CoT (to understand "how prompts elicit reasoning")One-line summary: architecture before scale, scale before alignment, applications last. We recommend reading Transformer Architecture as a foundation first, then returning to this checklist.
7. Three Companion Materials Per Paper
Reading papers alone has limited efficiency. For each paper on the list, pair it with three things: official code / illustrated blogs / this site's deep-dive pages. Read visuals first to build intuition, then read the original to check details, and finally run the code.
| # | Paper | Illustrated / Blog | Official Code or Implementation | This Site's Deep-Dive |
|---|---|---|---|---|
| 1 | Attention | The Illustrated Transformer | The Annotated Transformer | Deep-dive |
| 2 | GPT-1 | OpenAI Blog | Community implementations (Hugging Face repos) | Deep-dive |
| 3 | GPT-2 | OpenAI Blog | OpenAI/gpt-2 | Deep-dive |
| 4 | BERT | Jay Alammar Illustrated | google-research/bert | Deep-dive |
| 5 | GPT-3 | OpenAI Blog + third-party long reads | Commercial API, no open weights | Reading Paths Checklist |
| 6 | Scaling Laws | OpenAI Blog | No official code (reproducible experiments) | Deep-dive |
| 7 | Chinchilla | DeepMind Blog | No official code | Deep-dive |
| 8 | InstructGPT | OpenAI Blog | Commercial API, no open weights | Deep-dive |
| 9 | RLHF / DPO | Multiple illustrated guides (search "DPO explained") | Hugging Face TRL implementation | Deep-dive |
| 10 | LoRA | HF PEFT Docs | microsoft/LoRA | Deep-dive |
| 11 | FlashAttention | Paper blog + video walkthroughs | Dao-AILab/flash-attention | Deep-dive |
| 12 | CoT | Multiple Chinese-language walkthroughs | google-research/chain-of-thought | Deep-dive |
| 13 | RAG | HF Blog | facebookresearch/rag | Deep-dive |
Order of using materials
Illustrations (20 min) → Original paper deep-dive (1–2 hours) → Run code or read official implementation (optional) — don't do it in reverse: diving straight into code often leaves you lost in engineering details. Get the mechanism first, then look at the implementation.
8. An 8-Week Execution Plan for Path B (Example)
Path B is the most prone to "having a checklist but no action." Here's a copy-paste 8-week plan (4–5 hours per week):
| Week | Topic | Deep-Dive Papers | Practice Action | Deliverable |
|---|---|---|---|---|
| W1 | Architecture foundation | Attention (second pass) | Read Transformer Architecture | Three-sentence note |
| W2 | Pretraining paradigm | GPT-1, GPT-2 | Browse gpt-2 official code | Card ×1 |
| W3 | Scaling laws | Scaling Laws, Chinchilla | Estimate your data needs using 1:20 ratio | Ratio notes |
| W4 | Alignment | InstructGPT | Read alignment concept page | Card ×1 |
| W5 | Fine-tuning | LoRA | Run a LoRA fine-tuning demo (see Fine-Tuning Practice) | Transfer risk checklist |
| W6 | Retrieval | RAG | Build minimal RAG demo (see RAG in Practice) | Card ×1 |
| W7 | Reasoning | FlashAttention, CoT | Run a long-context / prompting comparison experiment | Evaluation notes |
| W8 | Wrap-up | Review 8 weeks + frontier selection | Update personal map + write summary | Personal paper map |
Flexibility in the plan
If W5/W6 demos don't have enough time, defer to Week 9 — paper reading rhythm can be interrupted, but deep-dive quality must not be compromised. It's better to finish 6 deep-dives in 8 weeks than to gulp down 20 abstracts in 8 weeks.
9. Time Budget and Reading Rhythm
"No time" is the biggest enemy of "reading papers." Here are three budget tiers:
| Available Time | Weekly Recommendation | What You Can Finish in a Month |
|---|---|---|
| 20 min/day | 2 abstract-level papers + 1 deep-dive on weekend | Path A + first 3 steps of Path B |
| 1 hour/day | 2 deep-dives + 1 card note | Full Path B + some frontier |
| 2+ hours/day | 3 deep-dives + 1 reproduction experiment | Path B + Path B start of Path C |
The "minimum viable action" when you're short on time
No matter how busy, keep one action alive: scan arXiv headlines, pick 1 paper to read the abstract from (5 min). This action maintains "trend sensitivity," and picking it back up after a break is hard. The full tracking method is in Section 6 of Reading Discipline & FAQ.
10. From Checklist to Deep-Dive: Common Questions
Q: The 13 papers on the checklist feel like too many — which should I deep-dive first? A: Look up by your scenario in Section 4. If you have no clear scenario, start with these four papers in order: 1 → 5 → 8 → 10. These four cover the minimal skeleton of four major threads: architecture, scale, alignment, applications.
Q: How long does a deep-dive take per paper? A: First pass: 15 min; second pass: 1–2 hours; third pass: half a day to a few days. For Path B, only the second pass + limitations analysis is needed.
Q: Reading the English original is tough. What should I do? A: It's fine to use translation tools for the first pass, but your notes must use the original English terminology (e.g., self-attention, ablation), otherwise you won't be able to cross-reference literature or answer interview questions. Term quick-reference: glossary.
Q: I forget everything after reading. What should I do? A: It's not a memory problem — it's a lack of outputs. Use the card note method to write a five-section card per paper, and file it into your personal paper map.
Common pitfall
Counting "50 abstracts read" as "having read papers." Abstract-level reading only builds trivia — not judgment. At least 5–8 papers on the checklist must reach "deep-dive + note" level, otherwise you'll freeze when an interviewer asks "what do the ablations prove?"
Further Reading
- Paper Map — A horizontal expansion of this checklist: place each paper on the timeline and topic coordinate system
- Core Paper Deep Dives — In-depth readings of 11 papers from this essential checklist (background / method / experiments / limitations)
- Frontier Trends — New directions beyond the checklist: o1, Mamba, MoE scaling, and more
- Reading Discipline & FAQ — Complete methodology: three-pass reading, note-taking, judging paper quality
- Evolution Timeline — A narrative version of the paper timeline, suitable for beginners as a companion read
References
Below are the official original links (arXiv) for the essential reading list papers — full text freely accessible:
- Vaswani et al. Attention Is All You Need (2017) — The Transformer origin paper
- Radford et al. Improving Language Understanding by Generative Pre-Training (2018) — GPT-1 (OpenAI official PDF; not formally published on arXiv)
- Radford et al. Language Models are Unsupervised Multitask Learners (2019) — GPT-2
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (2019) — BERT
- Brown et al. Language Models are Few-Shot Learners (2020) — GPT-3
- Kaplan et al. Scaling Laws for Neural Language Models (2020) — Scaling Laws
- Hoffmann et al. Training Compute-Optimal Large Language Models (2022) — Chinchilla
- Ouyang et al. Training Language Models to Follow Instructions with Human Feedback (2022) — InstructGPT
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (2021) — LoRA
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022) — FlashAttention
- Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) — CoT
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — RAG