Skip to content

Reading Paths

At a glance curated essential reading list from thousands of yearly arXiv papers: ~12 core papers, three paths (2-hour intro / engineering / research) with specific paper checklists and reading order, a "which sections to read" guide for each paper, and a quick-reference table for selecting papers by goal scenario.

Reading Paths ​

One-sentence summary: This page is a decision table for "what to read, in what order, and to what depth." It organizes scattered arXiv papers into three actionable paths, so whether you have two hours or plan to do research, you can find your starting point here.

1. Two Principles to Keep in Mind ​

1. Go Wide First, Then Deep ​

Read "broadly covering" milestone papers first to build your coordinate system (see Paper Map), then pick one or two for deep dives (see Core Paper Deep Dives). Jumping straight into a 2025 paper without context will leave you both confused and wasting the opportunity to build a global view.

Why does this order work? Because new papers assume you're "someone who knows the context": they assume you've read Transformer, know RLHF, and recognize MoE. Without a horizontal coordinate system, every sentence in a new paper references something you don't know. With a coordinate system, a new paper's incremental information is only a thin layer. The cost of reading a paper is "context," and context comes from horizontal accumulation.

2. Define "How Deep" for Each Paper ​

For the same paper, beginners read abstract + conclusion, engineers read method + experiments, and researchers read derivations + reproduction. Setting a depth target for each paper is more efficient than powering through the full text. The full depth tiering (the three-pass reading method) is in Reading Discipline & FAQ. Here's a quick heuristic:

Your GoalRecommended DepthCorresponding Pass
Understand trends, be able to chatAbstract + intro + conclusionFirst pass
Judge whether to apply to your workMethod + experiments + ablation + limitationsSecond pass
Propose improvements, do researchFull text + derivations + reproductionThird pass

2. Essential Reading List (~12 Papers) ​

The table below is the site's "essential core" covering the full pipeline from architecture to systems to alignment. Note: By spec count, this merges to ~12 papers; listed individually here it's 13 rows (RLHF and DPO are grouped as one entry). The deep-dive versions are on Core Paper Deep Dives.

#Paper (Year)Why It's Essential (One Sentence)Sections to Focus OnWhat You Can Answer After Reading
1Attention Is All You Need (2017)The origin of everything Transformer: self-attention + multi-head + positional encodingAbstract, §3 Model Architecture, §5 Experiments & ResultsWhy can self-attention replace RNN?
2GPT-1 (2018)First "generative pretraining + fine-tuning," opens the decoder routeAbstract, §3 Framework, Table 1 ResultsWhy does pretrain+fine-tune save labeled data?
3GPT-2 (2019)Proves "pretrained models can zero-shot transfer," introduces scale beliefAbstract, §2 Method, §4 ResultsHow does zero-shot happen "for free"?
4BERT (2019)Bidirectional encoder + masked language modeling, dominant paradigm for understanding tasksAbstract, §3 Pre-training Objectives, §4 AblationsWhy is bidirectional more important for understanding?
5GPT-3 (2020)175B parameters + few-shot, an empirical manifesto for scalingAbstract, §2 Method, §3.9 Limitations, §3 ResultsHow does few-shot performance scale with model size?
6Scaling Laws (2020)Loss decreases as a power law of params/data/compute, guides "where to spend money"Abstract, §3 Power Laws, §4 Transfer, Figures 1–4Should I add parameters or data?
7Chinchilla (2022)Compute-optimal ratio: params to tokens ≈ 1:20, corrects "bigger is always better"Abstract, §4 Optimal Ratio, Table 3 ComparisonWhy did a 70B model beat a 280B model?
8InstructGPT (2022)RLHF three-step method, the technical mother of ChatGPTAbstract, §3 Method (SFT/RM/PPO), §4 Human EvaluationHow do models go from "can talk" to "obeys instructions"?
9RLHF (2017) & DPO (2023)A pair: preference learning evolves from "training a reward model" to "using preferences directly"RLHF: §2 Method; DPO: §3 Derivation, Table 1Why can DPO skip the reward model?
10LoRA (2021)Parameter-efficient fine-tuning via low-rank updates, 10,000× reduction in fine-tuning costAbstract, §3.1 Low-Rank Parameterization, §4 Experiment TableWhy does the low-rank assumption hold?
11FlashAttention (2022)I/O-aware attention optimization, memory from O(n²) to linear, prerequisite for long contextAbstract, §3 Algorithm, §5 Speedup TableIs the attention bottleneck compute or memory bandwidth?
12Chain-of-Thought (2022)The key to LLM reasoning ability, a watershed for promptingAbstract, §2 Method, GSM8K Table, §4 AblationsWhy doesn't CoT work for small models?
13RAG (2020)The RAG paradigm origin, classic solution for hallucination mitigation and knowledge injectionAbstract, §3 Models (RAG-sequence/token), §4 ExperimentsWhy is an external "retrieval" better than "memorization"?

How to use this table

You don't need to read them in numbered order. Beginners: start with 1, 5, 8, 10; Engineers: focus on 10, 11, 13; Researchers: read 6, 7, 9 thoroughly. Internal order within each path is in Section 6's dependency diagram.

3. Three Paths: Specific Checklists and Reading Order ​

Path A: 2-Hour Quick Start ​

Goal: Build the skeleton of "how Transformer works + why LLMs are powerful + how to make them obedient," so you can talk about it and not get lost in blogs.

text
Step 1 (40 min): Attention Is All You Need → Read only abstract, intro, conclusion + view The Illustrated Transformer visuals
Step 2 (30 min): GPT-1 → GPT-2 abstract comparison, understand "decoder + pretraining" route
Step 3 (30 min): GPT-3 abstract + limitations section, understand "scale + few-shot"
Step 4 (20 min): InstructGPT abstract + Figure 1, understand RLHF's three steps
Deliverable: Can say one sentence each for "self-attention, causal mask, pretraining, few-shot, RLHF"

Requirement for this path: You don't need to understand formulas, but you must be able to answer three sentences per paper: "problem / method / effect." If a concept blocks you (e.g., causal mask), check Transformer Architecture or the glossary — don't struggle through it.

Path B: Engineering Implementation ​

Goal: Be able to judge "whether a paper's method can be applied to my work," and actually fine-tune or deploy it.

text
Prerequisite: Complete Path A first
Step 1: LoRA deep-dive method + experiments → cross-reference with Practice page [Fine-Tuning Practice: Full LoRA Workflow](/practice/fine-tuning-practice)
Step 2: RAG paper deep-dive §3–§4 → cross-reference with [case-studies RAG](/case-studies/rag) and [RAG in Practice](/practice/rag-in-practice)
Step 3: FlashAttention read abstract + algorithm → understand long context and inference speedup (paired with [Inference Fundamentals](/concepts/inference-fundamentals))
Step 4: Chinchilla read abstract + optimal ratio → understand "should I add data or parameters first?"
Step 5: CoT read method + GSM8K results → prompting engineering basis (paired with [Prompting](/concepts/prompting))
Deliverable: Can select "fine-tune vs. RAG vs. prompting" for a project and justify with paper evidence

Requirement for this path: For each paper, read "ablation + hyperparameters + limitations" and write a "risk checklist for transfer to my scenario." After finishing each paper, we recommend also checking the corresponding practice page in the practice guide — papers give principles, practice pages give the gotchas.

Path C: Research / Frontier Tracking ​

Goal: Stand at the 2025 frontier, read the latest arXiv, propose and validate improvements.

text
Prerequisite: Complete Paths A + B
Step 1: Read Scaling Laws full text (including appendix) and Chinchilla §4 derivation
Step 2: RLHF original paper (Christiano 2017) → InstructGPT → DPO, progressive reading to understand the evolution of preference learning
Step 3: Pick your own direction (alignment / systems / applications) and trace citation chains backwards from [Frontier Trends](/papers/frontier)
Step 4: Reproduce one paper (recommend a minimal implementation of LoRA or FlashAttention), build a "read + write" loop
Deliverable: Can write a one-page related work and propose a verifiable improvement hypothesis

Requirement for this path: Reading papers isn't the end goal — writing is. Every third-pass deep-dive must produce: a mechanism explanation, a list of untested experiments, and an improvement hypothesis. Even if you don't validate the hypothesis, write it down — in three months you'll be surprised at how far you've come.

Comparison of the Three Paths ​

DimensionPath A: IntroPath B: EngineeringPath C: Research
Target AudienceBeginners, product / non-ML rolesML engineers, application devsGrad students, researchers
Total Time2 hours2–3 weeks (1 hour/day)Ongoing
Deep-Dive Count4 papers at abstract level6–8 papers deep15+ papers + continuous tracking
Key ActionsView visuals, grasp conceptsReproduce experiments, write assessmentsDerivations, reproduction, propose hypotheses
DeliverablesThree-sentence notesTransfer risk checklistRelated work + improvement hypotheses
Validation CriteriaCan answer "why the GPT series succeeded"Can deploy a feature based on a paper methodCan write related work, can defend your work
Corresponding PagesCore Paper Deep DivesDeep-dives + Frontier Trends + practice modulesPaper Map + Frontier Trends

4. Paper Selection Quick-Reference by Goal Scenario ​

Not sure where to start? Look up by your current goal:

Your ScenarioRead These FirstPaired Pages
Want to understand Transformer and attention1, 11Transformer Architecture
Want to understand "why bigger models are stronger"6, 7, 5Scaling Laws
Want to do fine-tuning (SFT / LoRA)10, 2Fine-Tuning: SFT and PEFT, Fine-Tuning Practice
Want to understand alignment and RLHF / DPO8, 9Alignment: RLHF and DPO
Want to do RAG or reduce hallucination13, 12RAG Case Studies, RAG in Practice
Want to optimize inference / deployment performance11, 3Inference Fundamentals, Deployment Practice
Want to do evaluation / benchmarks6, 7Evaluation and Benchmarks, Evals in Practice
Preparing for ML interviews1, 5, 8, 10, 11, 13 deepInterview Question Bank

5. Generic Template: "Which Sections to Read in Each Paper" ​

Paper structures vary, but most follow an IMRaD structure (Introduction–Method–Results–Discussion). A generic trade-off and selection table:

SectionBeginner ReaderEngineer ReaderResearcher Reader
Abstract + figure summariesRequiredRequiredRequired
IntroductionRequiredRequiredRequired (focus on "gap" argument)
Related WorkCan skipSkimRequired (follow citation chains)
MethodFigure onlyDeep-dive + formulasDeep-dive + derivations
ExperimentsMain table onlyDeep-dive ablations + hyperparamsDeep-dive + compare against reproduction
Limitations / DiscussionRequiredRequired (judge transferability)Required (find improvement points)

Specific selections for the essential reading list are already in the "reading approach" column of the Section 2 table. The full three-pass reading flow is in Reading Discipline & FAQ.

Why "Limitations" is required for everyone

Many readers skip limitations — and miss the most valuable information. The limitations section typically contains: under what conditions the method fails, what experiments the authors didn't do, and hints for future work — these are respectively the boundaries for engineering transfer, sources of improvement ideas, and clues for trend judgment.

There's a clear inheritance chain between papers — reading in dependency order pays dividends:

text
Attention Is All You Need ─┬─→ GPT-1 → GPT-2 → GPT-3 ─┬─→ InstructGPT → DPO
                            └─→ BERT ──────────────────┘
                                                             ↓
Scaling Laws ←─ Chinchilla ←─ InstructGPT (to understand "the cost of alignment")
Attention ──→ FlashAttention (to understand "how to speed up the same architecture")
GPT-3 ──→ RAG (to understand "parametric knowledge vs. retrieved knowledge")
GPT-3 ──→ CoT (to understand "how prompts elicit reasoning")

One-line summary: architecture before scale, scale before alignment, applications last. We recommend reading Transformer Architecture as a foundation first, then returning to this checklist.

7. Three Companion Materials Per Paper ​

Reading papers alone has limited efficiency. For each paper on the list, pair it with three things: official code / illustrated blogs / this site's deep-dive pages. Read visuals first to build intuition, then read the original to check details, and finally run the code.

#PaperIllustrated / BlogOfficial Code or ImplementationThis Site's Deep-Dive
1AttentionThe Illustrated TransformerThe Annotated TransformerDeep-dive
2GPT-1OpenAI BlogCommunity implementations (Hugging Face repos)Deep-dive
3GPT-2OpenAI BlogOpenAI/gpt-2Deep-dive
4BERTJay Alammar Illustratedgoogle-research/bertDeep-dive
5GPT-3OpenAI Blog + third-party long readsCommercial API, no open weightsReading Paths Checklist
6Scaling LawsOpenAI BlogNo official code (reproducible experiments)Deep-dive
7ChinchillaDeepMind BlogNo official codeDeep-dive
8InstructGPTOpenAI BlogCommercial API, no open weightsDeep-dive
9RLHF / DPOMultiple illustrated guides (search "DPO explained")Hugging Face TRL implementationDeep-dive
10LoRAHF PEFT Docsmicrosoft/LoRADeep-dive
11FlashAttentionPaper blog + video walkthroughsDao-AILab/flash-attentionDeep-dive
12CoTMultiple Chinese-language walkthroughsgoogle-research/chain-of-thoughtDeep-dive
13RAGHF Blogfacebookresearch/ragDeep-dive

Order of using materials

Illustrations (20 min) → Original paper deep-dive (1–2 hours) → Run code or read official implementation (optional) — don't do it in reverse: diving straight into code often leaves you lost in engineering details. Get the mechanism first, then look at the implementation.

8. An 8-Week Execution Plan for Path B (Example) ​

Path B is the most prone to "having a checklist but no action." Here's a copy-paste 8-week plan (4–5 hours per week):

WeekTopicDeep-Dive PapersPractice ActionDeliverable
W1Architecture foundationAttention (second pass)Read Transformer ArchitectureThree-sentence note
W2Pretraining paradigmGPT-1, GPT-2Browse gpt-2 official codeCard ×1
W3Scaling lawsScaling Laws, ChinchillaEstimate your data needs using 1:20 ratioRatio notes
W4AlignmentInstructGPTRead alignment concept pageCard ×1
W5Fine-tuningLoRARun a LoRA fine-tuning demo (see Fine-Tuning Practice)Transfer risk checklist
W6RetrievalRAGBuild minimal RAG demo (see RAG in Practice)Card ×1
W7ReasoningFlashAttention, CoTRun a long-context / prompting comparison experimentEvaluation notes
W8Wrap-upReview 8 weeks + frontier selectionUpdate personal map + write summaryPersonal paper map

Flexibility in the plan

If W5/W6 demos don't have enough time, defer to Week 9 — paper reading rhythm can be interrupted, but deep-dive quality must not be compromised. It's better to finish 6 deep-dives in 8 weeks than to gulp down 20 abstracts in 8 weeks.

9. Time Budget and Reading Rhythm ​

"No time" is the biggest enemy of "reading papers." Here are three budget tiers:

Available TimeWeekly RecommendationWhat You Can Finish in a Month
20 min/day2 abstract-level papers + 1 deep-dive on weekendPath A + first 3 steps of Path B
1 hour/day2 deep-dives + 1 card noteFull Path B + some frontier
2+ hours/day3 deep-dives + 1 reproduction experimentPath B + Path B start of Path C

The "minimum viable action" when you're short on time

No matter how busy, keep one action alive: scan arXiv headlines, pick 1 paper to read the abstract from (5 min). This action maintains "trend sensitivity," and picking it back up after a break is hard. The full tracking method is in Section 6 of Reading Discipline & FAQ.

10. From Checklist to Deep-Dive: Common Questions ​

Q: The 13 papers on the checklist feel like too many — which should I deep-dive first? A: Look up by your scenario in Section 4. If you have no clear scenario, start with these four papers in order: 1 → 5 → 8 → 10. These four cover the minimal skeleton of four major threads: architecture, scale, alignment, applications.

Q: How long does a deep-dive take per paper? A: First pass: 15 min; second pass: 1–2 hours; third pass: half a day to a few days. For Path B, only the second pass + limitations analysis is needed.

Q: Reading the English original is tough. What should I do? A: It's fine to use translation tools for the first pass, but your notes must use the original English terminology (e.g., self-attention, ablation), otherwise you won't be able to cross-reference literature or answer interview questions. Term quick-reference: glossary.

Q: I forget everything after reading. What should I do? A: It's not a memory problem — it's a lack of outputs. Use the card note method to write a five-section card per paper, and file it into your personal paper map.

Common pitfall

Counting "50 abstracts read" as "having read papers." Abstract-level reading only builds trivia — not judgment. At least 5–8 papers on the checklist must reach "deep-dive + note" level, otherwise you'll freeze when an interviewer asks "what do the ablations prove?"

Further Reading ​

  • Paper Map — A horizontal expansion of this checklist: place each paper on the timeline and topic coordinate system
  • Core Paper Deep Dives — In-depth readings of 11 papers from this essential checklist (background / method / experiments / limitations)
  • Frontier Trends — New directions beyond the checklist: o1, Mamba, MoE scaling, and more
  • Reading Discipline & FAQ — Complete methodology: three-pass reading, note-taking, judging paper quality
  • Evolution Timeline — A narrative version of the paper timeline, suitable for beginners as a companion read

References ​

Below are the official original links (arXiv) for the essential reading list papers — full text freely accessible: