Theme
Getting Started
One-sentence summary: This paper section takes you from "can call an API" to "can read the latest arXiv paper from 2025". It doesn't aim to devour every paper cover-to-cover all at once — instead, it first answers three questions: "what to read, how to read it, and to what depth."
1. Why LLM Practitioners Must Read Papers
1. LLMs Are a "Paper-Driven" Field, with a Technology Half-Life of Only a Few Months
Every API, every tutorial you use today has a paper behind it: gpt-4o is backed by native multimodal architecture choices, DeepSeek-R1 by reinforcement learning and inference-time scaling, and vLLM by PagedAttention. "The latest" technology from three years ago is now textbook material. Papers are the primary source of this information — new methods always appear first as preprints on arXiv, and only months later show up in blogs, courses, and framework documentation.
People who don't read papers only get information that is filtered, exaggerated, or even distorted through secondhand retellings. People who read papers can directly assess "under what data, compute, and evaluation protocols was this SOTA result achieved." The former chase trends; the latter see through them.
2. Secondhand Information Distorts: Blogs Handle Emotion, Papers Handle Facts
Information pollution in the LLM space is worse than in traditional ML — marketing copy, social media posts, and investment bank reports all talk about "large models." A "groundbreaking breakthrough" from 2023 might just be a random fluctuation on a benchmark by 2025. Reading papers directly helps you build anchors in reality:
| Question | Reading Only Blogs | Reading Papers |
|---|---|---|
| "Model A beats Model B" | Accept the conclusion | Probe the benchmark, few-shot setup, and data contamination |
| "This method improves by 30%" | Get excited | Check whether the improvement comes from ablation or a full replacement |
| "Large models can think" | Share it | Go back to the evidence: what evaluation tasks, who evaluated them? |
| "Training requires 10,000 GPUs" | Get discouraged | See whether it's the paper's ideal config or a reproducible one |
3. Understand the Principles to Transfer Them to Your Own Work
Tutorials teach you model.generate() — papers tell you why that interface was designed the way it is: Why does self-attention divide by √d_k? Why do we use KV Cache in the decode phase? Why does LoRA constrain update matrices to low-rank? Why does positional encoding use length extrapolation rather than hard truncation? Every design corresponds to an ablation, an experiment, or a failure in the original paper. What you get from papers is not "how to use it" but "why it was designed that way" — that's the watershed moment between "someone who can use frameworks" and "someone who can design systems."
The site's core knowledge modules (such as Transformer Architecture, Scaling Laws, Alignment: RLHF and DPO) give you a concept map. This paper section fills in why each place on that map has the name it does.
4. Hard Career Skills: Interviews, Promotions, and Judgment
In LLM job interviews (see the interview question bank), "discussing a paper in depth" is almost guaranteed to come up. Interviewers aren't asking for the abstract — they want to know:
- What problem does this paper solve? Why does it matter?
- Why didn't baselines work? What is the core contribution?
- What do the ablation studies prove? What's novel vs. what's just engineering tuning?
- In your scenario, where would this method fail?
Only someone who has carefully read the original paper can answer these. What interviewers really want to see is "you have the ability to read papers" — which is the core skill of everyday R&D. Promotion defense, tech selection reviews, and team tech talks are all different expressions of the same ability.
5. Three Information Sources Have Radically Different Roles
| Information Source | Role | Information Density | Timeliness | Distortion Risk | Best For |
|---|---|---|---|---|---|
| Tutorials / Courses | Help you get started, build intuition | Low | Slow (often 1+ year behind) | Medium | Beginners |
| Blogs / Social Media | Track trends, relay conclusions | Medium | Fast but often exaggerated in headlines | High | Beginners + trend chasers |
| Original Papers | The only anchor for facts and mechanisms | High | Earliest (preprints release immediately) | Low | Practitioners (essential) |
Takeaway
Tutorials handle "first meeting," blogs handle "staying aware," and papers handle "building judgment." They're not substitutes — they're divisions of labor. People who only read the first two are forever swimming in someone else's conclusions.
2. Three Levels of Paper Reading: Know, Understand, Apply
Many people "have read a lot of papers" but haven't improved, because they're stuck at the first level. After reading a paper, ask yourself which level you're at:
| Level | Signature Output | Test Question | Next Step |
|---|---|---|---|
| Know | Can restate the abstract | "What is this paper about?" | Explain it to someone else |
| Understand | Can explain the mechanism | "Why does it work? What do ablations prove?" | Write a card note (see Reading Discipline & FAQ) |
| Can Apply | Can transfer to your scenario | "What happens with my data, compute, or eval?" | Write a minimal experiment or code reproduction |
The three levels correspond to increasing reading depth, and to different "how deep" levels for the three paths in Reading Paths. Most people are stuck at "Know" — they've bookmarked 200 papers and can name their titles, but can't explain any single one. Every method here (the three-pass reading method, card notes, judgment criteria) is designed to push you toward "Understand" and "Can Apply."
3. 12 Terms You Must Know to Read Papers
Papers have their own jargon. A single abstract full of unfamiliar terms can be enough to scare you off. Here are the 12 most common — scan them before you start reading:
| Term | Meaning | Common Location |
|---|---|---|
| Abstract | Summary: problem + method + conclusion condensed | Beginning |
| Motivation | Why the authors did this work | Introduction |
| Contribution | What's new in this paper (usually 3–4 bullet points) | End of Introduction |
| Method / Approach | Core design, formulas, algorithms | Middle section |
| Ablation | Deconstructing a component to measure its contribution | Experiments |
| Baseline | Existing methods used for comparison | Experiments |
| SOTA (State-of-the-Art) | Best performance at the time of writing | Experiments / Abstract |
| Setting / Setup | Experimental config: data, hyperparams, compute, eval protocol | Start of Experiments |
| Evaluation Metric | BLEU, Perplexity, Accuracy, etc. | Experiments |
| Limitations | Self-identified weaknesses | Before Conclusion |
| Discussion | Interpretation of results, relationship to other work | Before Conclusion |
| Reproducibility | Code/data availability | End of paper / Appendix |
| Pretrained / Fine-tune | Pretrain on general corpus, then adapt on a specific task | Introduction / Method |
| Zero-shot / Few-shot | No examples / a few examples for in-context reasoning | Experiments |
| Inference / Decoding | Post-training generation phase (sampling, KV Cache) | Method / System sections |
| Scaling | Growth patterns of parameters, data, and compute | Throughout (esp. Conclusion) |
| Benchmark | Standard task suites for comparison (e.g., MMLU, GSM8K) | Experiments |
| Contamination | Test data leaked into training data, inflating scores | Discussion / Appendix |
Don't memorize terms — read three papers and you'll know them
A full glossary is available in the Resources section. The key takeaway: use "term-based navigation" when reading papers — when you see Abstract, ask "what's the problem?"; when you see Ablation, ask "which component is the most valuable?"; when you see Limitations, ask "when does this method fail?"
4. Where to Start: Three Paths
There's no one-size-fits-all depth for reading papers. Choose a path based on your goal. For detailed checklists and reading order, see Reading Paths. Here's a quick overview:
Path A: 2-Hour Quick Start
├─ Read only abstract + intro + conclusion + key diagrams
├─ Supplement with one deep-dive note or illustrated blog
└─ Deliverable: Can clearly state "what problem it solves, what the method is in one sentence, what the result was"
Path B: Engineering Implementation
├─ Deep-dive: method + experiments + ablations
├─ Focus on hyperparams, limitations, and discussion
└─ Deliverable: Can judge "whether this method can be applied to my business data"
Path C: Research / Frontier Tracking
├─ Deep-dive everything: full text + math + reproduce experiments
├─ Read references backwards, compare subsequent work
└─ Deliverable: Can propose and validate your own improvements| Path | Who It's For | Reading Depth | Total Time | Where to Start |
|---|---|---|---|---|
| 2-Hour Quick Start | Learners who just finished the basics | Abstract + intro + conclusion + diagrams | 2 hours | Beginner articles in Core Paper Deep Dives |
| Engineering Implementation | Engineers who want to solve real problems | Full text + ablations + limitations | 2–3 weeks | Core Paper Deep Dives + Frontier Trends |
| Research / Frontier | Researchers, grad students, trend trackers | Full text + derivations + reproduction | Ongoing | Paper Map to build your coordinate system, then go deeper |
Path selection advice:
- Still building foundations (haven't finished the concepts section yet): Start with Path A. The priority is "building a concept skeleton" — don't rush.
- Need to ship to production / preparing for interviews: Go with Path B. The priority is "being able to explain 5–8 papers in depth + transfer them to your work." Path A is just warm-up.
- Want to publish / do research: Path C is your long-term state; Paths A and B are its daily building blocks.
One sentence for beginners
Always pick "classic and readable" for your first paper, not "newest and hottest." Classic papers have stood the test of time and are cited heavily by subsequent work — they're usually written more clearly than papers rushed through peer review. Build the confidence that "I can read a whole paper" before worrying about quantity.
5. What's in This Section
The paper section has six pages. We recommend visiting them in this order:
| Page | One-Sentence Description | What You'll Get |
|---|---|---|
| Reading Paths | The section's entry point: essential reading list + three paths with specific papers and order | A paper-reading roadmap organized by goal |
| Paper Map | Key papers from decades of LLM research arranged on a map by timeline + topic | Macro coordinate system: which papers are origins, which are milestones |
| Core Paper Deep Dives | In-depth reading of 11 field-defining papers: background, method, experiments, limitations | The ability to truly understand principles, plus interview talking points |
| Frontier Trends | Nine major trends from 2023–2025: o1, long context, MoE, Mamba, … | Stay current, know which way the wind is blowing |
| Reading Discipline & FAQ | Methodology: three-pass reading, note-taking, judging paper quality, following arXiv | Sustainable reading habits, not a one-time burst of enthusiasm |
The logic between these pages is map first, then deep-dives, then frontier: the map answers "what to read," deep-dives answer "how deep to read," and frontier answers "where to go next." If you're not comfortable with terminology, refer to the glossary anytime. If you want domain context first, check out the evolution timeline.
6. Common Mistakes: Five Ways to Waste Your Time
| Mistake | Consequence | Correct Approach |
|---|---|---|
| Reading only the abstract, then sharing | Accumulates trivia, zero judgment | At least finish the first pass of three-pass reading and answer three sentences |
| Replacing original papers with blogs | Led around by secondhand conclusions | Blogs build intuition, originals verify judgment — read blogs first, then originals |
| "Bookmarked" equals "read" | Your bookmark folder becomes a graveyard | Every paper must produce an output (note / retelling / experiment) |
| Reading linearly from the first formula | Stuck in Section 3 after two hours | Start with abstract + diagrams, then decide how deep to go (a paper is a reference manual, not a novel) |
| Chasing only new papers without reading classics | No coordinate system; new papers are also unreadable | Build foundations with classics (see essential reading list), then follow the frontier |
7. How to Schedule Time Each Week
You don't need large blocks of time at once. A fixed 2–3 hours per week is more effective than cramming 10 hours:
| Time Slot | Content | Duration |
|---|---|---|
| Monday, 15 min | Scan arXiv / Hugging Face Papers headlines (see Following arXiv) | 15 min |
| Wednesday, 45 min | Deep-dive the week's one paper (passes 1 and 2 of three-pass reading) | 45 min |
| Saturday, 1 hour | Finish pass 3 + write card notes | 1 hour |
| Sunday, 30 min | File notes, update your personal paper map | 30 min |
Rhythm matters more than intensity. After 8 weeks, you'll have 8 card notes, a clear narrative thread, and an intuition for "which papers are worth reading."
8. 10 Questions to Ask Before Reading Any Paper
Use these 10 questions as a template before you open any paper. They'll quickly help you judge "whether to deep-dive and where to focus":
- What problem is it trying to solve? Is the problem real, or was it contrived just to justify the method?
- Why didn't previous baselines work? Without an answer to this, the method is likely self-indulgent.
- What's the core innovation? Can you state it in one sentence? If you can't, the authors haven't figured it out either.
- What assumptions does it make? Under what conditions do those assumptions hold? Data, compute, task distribution.
- How are experiments set up? Data, hyperparams, compute, evaluation protocol — miss one and you can't trust the conclusions.
- Does the improvement come from innovation or brute force? Check ablations: remove the novel component, see how much performance drops.
- What are the authors' stated limitations? Papers that thoughtfully write Limitations tend to be more credible.
- What experiments are missing? This is where improvement opportunities lie — and also interview bonus points.
- In my scenario, where is this most likely to fail? The core question for engineering deployment.
- What reusable things does it leave behind? Code, data, models, or just an intuition?
Why "pre-reading" questions?
You should write these down before opening the full text. Having your own questions before reading prevents you from passively following the authors' narrative — reading a paper is fundamentally a "conversation," not a "lecture."
9. Minimal Action Checklist
After reading this page, follow this checklist to get started in your first week:
text
□ Choose a path (A/B/C) and tell a friend your choice
□ Open [Reading Paths](/papers/paths) and note down your first paper
□ Open [Core Paper Deep Dives](/papers/core-papers) and read one section
□ Use "term-based navigation" to scan that paper's title, abstract, and figures
□ Complete the first pass of [three-pass reading](/papers/faq), write down three sentences
□ Set up your notes file (or Obsidian vault) and write your first card
□ Subscribe to arXiv's cs.CL / cs.LG (see [Following arXiv](/papers/faq))10. The Most Important Piece of Advice
Take away mechanisms, not numbers
The numbers in papers are often the most memorable: "GPT-3 has 175B parameters," "GSM8K accuracy went from 18% to 57%." But remember:
Every absolute score depends on the data, compute, preprocessing, and evaluation protocol at the time. Change the data distribution, the hardware budget, or the evaluation script — and those numbers may no longer hold.
Three things are truly worth taking away:
- Mechanisms — Why does this method work? What structural problem does it solve? (e.g., FlashAttention solves the memory wall at the IO level for attention)
- Failure analysis — What experiments were missing? What approaches failed? What limitations did the authors state? Failures are often more valuable than successes.
- Boundary conditions — Under what conditions does this method hold? When does it fail? (e.g., CoT doesn't improve small models; it only emerges in sufficiently large models)
Leaderboard numbers go stale. Mechanisms and boundaries don't. This "read mechanisms, read boundaries" discipline is fully laid out in Reading Discipline & FAQ.
Further Reading
- Reading Paths — The next step from here: choose your first path and get a specific paper checklist
- What Is a Large Language Model — Domain concept foundation; we recommend building this global picture before reading papers
- Paper Map — Want to build a macro coordinate system first? Start here
- Glossary — Quick-term reference for paper reading
- Evolution Timeline — A popular version of the paper timeline, cross-referenced with the map
References
All of these are real, publicly available resources for self-directed deepening:
- arXiv Preprint Server — The primary publishing venue for most LLM papers; full text freely available
- Papers with Code — Paper + code + benchmark aggregation; useful for verifying reproducibility and SOTA claims
- Karpathy. State of GPT (2023 lecture) — A lecture-style walkthrough of "how large models are trained," bridging papers and engineering
- Jay Alammar. The Illustrated Transformer (2018) — An illustrated version of Attention Is All You Need — a model example of reading papers visually
- The Annotated Transformer (Harvard NLP) — A deep-dive Transformer paper with line-by-line annotations and runnable code