Skip to content

Getting Started

At a glance Why LLM practitioners must read papers. This page clarifies the three layers of value in paper reading, outlines three paths (2-hour intro, engineering implementation, and research), introduces the six-page paper section map, lists common pitfalls and time schedules, and shares one crucial reading recommendation.

Getting Started ​

One-sentence summary: This paper section takes you from "can call an API" to "can read the latest arXiv paper from 2025". It doesn't aim to devour every paper cover-to-cover all at once — instead, it first answers three questions: "what to read, how to read it, and to what depth."

1. Why LLM Practitioners Must Read Papers ​

1. LLMs Are a "Paper-Driven" Field, with a Technology Half-Life of Only a Few Months ​

Every API, every tutorial you use today has a paper behind it: gpt-4o is backed by native multimodal architecture choices, DeepSeek-R1 by reinforcement learning and inference-time scaling, and vLLM by PagedAttention. "The latest" technology from three years ago is now textbook material. Papers are the primary source of this information — new methods always appear first as preprints on arXiv, and only months later show up in blogs, courses, and framework documentation.

People who don't read papers only get information that is filtered, exaggerated, or even distorted through secondhand retellings. People who read papers can directly assess "under what data, compute, and evaluation protocols was this SOTA result achieved." The former chase trends; the latter see through them.

2. Secondhand Information Distorts: Blogs Handle Emotion, Papers Handle Facts ​

Information pollution in the LLM space is worse than in traditional ML — marketing copy, social media posts, and investment bank reports all talk about "large models." A "groundbreaking breakthrough" from 2023 might just be a random fluctuation on a benchmark by 2025. Reading papers directly helps you build anchors in reality:

QuestionReading Only BlogsReading Papers
"Model A beats Model B"Accept the conclusionProbe the benchmark, few-shot setup, and data contamination
"This method improves by 30%"Get excitedCheck whether the improvement comes from ablation or a full replacement
"Large models can think"Share itGo back to the evidence: what evaluation tasks, who evaluated them?
"Training requires 10,000 GPUs"Get discouragedSee whether it's the paper's ideal config or a reproducible one

3. Understand the Principles to Transfer Them to Your Own Work ​

Tutorials teach you model.generate() — papers tell you why that interface was designed the way it is: Why does self-attention divide by √d_k? Why do we use KV Cache in the decode phase? Why does LoRA constrain update matrices to low-rank? Why does positional encoding use length extrapolation rather than hard truncation? Every design corresponds to an ablation, an experiment, or a failure in the original paper. What you get from papers is not "how to use it" but "why it was designed that way" — that's the watershed moment between "someone who can use frameworks" and "someone who can design systems."

The site's core knowledge modules (such as Transformer Architecture, Scaling Laws, Alignment: RLHF and DPO) give you a concept map. This paper section fills in why each place on that map has the name it does.

4. Hard Career Skills: Interviews, Promotions, and Judgment ​

In LLM job interviews (see the interview question bank), "discussing a paper in depth" is almost guaranteed to come up. Interviewers aren't asking for the abstract — they want to know:

  • What problem does this paper solve? Why does it matter?
  • Why didn't baselines work? What is the core contribution?
  • What do the ablation studies prove? What's novel vs. what's just engineering tuning?
  • In your scenario, where would this method fail?

Only someone who has carefully read the original paper can answer these. What interviewers really want to see is "you have the ability to read papers" — which is the core skill of everyday R&D. Promotion defense, tech selection reviews, and team tech talks are all different expressions of the same ability.

5. Three Information Sources Have Radically Different Roles ​

Information SourceRoleInformation DensityTimelinessDistortion RiskBest For
Tutorials / CoursesHelp you get started, build intuitionLowSlow (often 1+ year behind)MediumBeginners
Blogs / Social MediaTrack trends, relay conclusionsMediumFast but often exaggerated in headlinesHighBeginners + trend chasers
Original PapersThe only anchor for facts and mechanismsHighEarliest (preprints release immediately)LowPractitioners (essential)

Takeaway

Tutorials handle "first meeting," blogs handle "staying aware," and papers handle "building judgment." They're not substitutes — they're divisions of labor. People who only read the first two are forever swimming in someone else's conclusions.

2. Three Levels of Paper Reading: Know, Understand, Apply ​

Many people "have read a lot of papers" but haven't improved, because they're stuck at the first level. After reading a paper, ask yourself which level you're at:

LevelSignature OutputTest QuestionNext Step
KnowCan restate the abstract"What is this paper about?"Explain it to someone else
UnderstandCan explain the mechanism"Why does it work? What do ablations prove?"Write a card note (see Reading Discipline & FAQ)
Can ApplyCan transfer to your scenario"What happens with my data, compute, or eval?"Write a minimal experiment or code reproduction

The three levels correspond to increasing reading depth, and to different "how deep" levels for the three paths in Reading Paths. Most people are stuck at "Know" — they've bookmarked 200 papers and can name their titles, but can't explain any single one. Every method here (the three-pass reading method, card notes, judgment criteria) is designed to push you toward "Understand" and "Can Apply."

3. 12 Terms You Must Know to Read Papers ​

Papers have their own jargon. A single abstract full of unfamiliar terms can be enough to scare you off. Here are the 12 most common — scan them before you start reading:

TermMeaningCommon Location
AbstractSummary: problem + method + conclusion condensedBeginning
MotivationWhy the authors did this workIntroduction
ContributionWhat's new in this paper (usually 3–4 bullet points)End of Introduction
Method / ApproachCore design, formulas, algorithmsMiddle section
AblationDeconstructing a component to measure its contributionExperiments
BaselineExisting methods used for comparisonExperiments
SOTA (State-of-the-Art)Best performance at the time of writingExperiments / Abstract
Setting / SetupExperimental config: data, hyperparams, compute, eval protocolStart of Experiments
Evaluation MetricBLEU, Perplexity, Accuracy, etc.Experiments
LimitationsSelf-identified weaknessesBefore Conclusion
DiscussionInterpretation of results, relationship to other workBefore Conclusion
ReproducibilityCode/data availabilityEnd of paper / Appendix
Pretrained / Fine-tunePretrain on general corpus, then adapt on a specific taskIntroduction / Method
Zero-shot / Few-shotNo examples / a few examples for in-context reasoningExperiments
Inference / DecodingPost-training generation phase (sampling, KV Cache)Method / System sections
ScalingGrowth patterns of parameters, data, and computeThroughout (esp. Conclusion)
BenchmarkStandard task suites for comparison (e.g., MMLU, GSM8K)Experiments
ContaminationTest data leaked into training data, inflating scoresDiscussion / Appendix

Don't memorize terms — read three papers and you'll know them

A full glossary is available in the Resources section. The key takeaway: use "term-based navigation" when reading papers — when you see Abstract, ask "what's the problem?"; when you see Ablation, ask "which component is the most valuable?"; when you see Limitations, ask "when does this method fail?"

4. Where to Start: Three Paths ​

There's no one-size-fits-all depth for reading papers. Choose a path based on your goal. For detailed checklists and reading order, see Reading Paths. Here's a quick overview:

Path A: 2-Hour Quick Start
├─ Read only abstract + intro + conclusion + key diagrams
├─ Supplement with one deep-dive note or illustrated blog
└─ Deliverable: Can clearly state "what problem it solves, what the method is in one sentence, what the result was"

Path B: Engineering Implementation
├─ Deep-dive: method + experiments + ablations
├─ Focus on hyperparams, limitations, and discussion
└─ Deliverable: Can judge "whether this method can be applied to my business data"

Path C: Research / Frontier Tracking
├─ Deep-dive everything: full text + math + reproduce experiments
├─ Read references backwards, compare subsequent work
└─ Deliverable: Can propose and validate your own improvements
PathWho It's ForReading DepthTotal TimeWhere to Start
2-Hour Quick StartLearners who just finished the basicsAbstract + intro + conclusion + diagrams2 hoursBeginner articles in Core Paper Deep Dives
Engineering ImplementationEngineers who want to solve real problemsFull text + ablations + limitations2–3 weeksCore Paper Deep Dives + Frontier Trends
Research / FrontierResearchers, grad students, trend trackersFull text + derivations + reproductionOngoingPaper Map to build your coordinate system, then go deeper

Path selection advice:

  • Still building foundations (haven't finished the concepts section yet): Start with Path A. The priority is "building a concept skeleton" — don't rush.
  • Need to ship to production / preparing for interviews: Go with Path B. The priority is "being able to explain 5–8 papers in depth + transfer them to your work." Path A is just warm-up.
  • Want to publish / do research: Path C is your long-term state; Paths A and B are its daily building blocks.

One sentence for beginners

Always pick "classic and readable" for your first paper, not "newest and hottest." Classic papers have stood the test of time and are cited heavily by subsequent work — they're usually written more clearly than papers rushed through peer review. Build the confidence that "I can read a whole paper" before worrying about quantity.

5. What's in This Section ​

The paper section has six pages. We recommend visiting them in this order:

PageOne-Sentence DescriptionWhat You'll Get
Reading PathsThe section's entry point: essential reading list + three paths with specific papers and orderA paper-reading roadmap organized by goal
Paper MapKey papers from decades of LLM research arranged on a map by timeline + topicMacro coordinate system: which papers are origins, which are milestones
Core Paper Deep DivesIn-depth reading of 11 field-defining papers: background, method, experiments, limitationsThe ability to truly understand principles, plus interview talking points
Frontier TrendsNine major trends from 2023–2025: o1, long context, MoE, Mamba, …Stay current, know which way the wind is blowing
Reading Discipline & FAQMethodology: three-pass reading, note-taking, judging paper quality, following arXivSustainable reading habits, not a one-time burst of enthusiasm

The logic between these pages is map first, then deep-dives, then frontier: the map answers "what to read," deep-dives answer "how deep to read," and frontier answers "where to go next." If you're not comfortable with terminology, refer to the glossary anytime. If you want domain context first, check out the evolution timeline.

6. Common Mistakes: Five Ways to Waste Your Time ​

MistakeConsequenceCorrect Approach
Reading only the abstract, then sharingAccumulates trivia, zero judgmentAt least finish the first pass of three-pass reading and answer three sentences
Replacing original papers with blogsLed around by secondhand conclusionsBlogs build intuition, originals verify judgment — read blogs first, then originals
"Bookmarked" equals "read"Your bookmark folder becomes a graveyardEvery paper must produce an output (note / retelling / experiment)
Reading linearly from the first formulaStuck in Section 3 after two hoursStart with abstract + diagrams, then decide how deep to go (a paper is a reference manual, not a novel)
Chasing only new papers without reading classicsNo coordinate system; new papers are also unreadableBuild foundations with classics (see essential reading list), then follow the frontier

7. How to Schedule Time Each Week ​

You don't need large blocks of time at once. A fixed 2–3 hours per week is more effective than cramming 10 hours:

Time SlotContentDuration
Monday, 15 minScan arXiv / Hugging Face Papers headlines (see Following arXiv)15 min
Wednesday, 45 minDeep-dive the week's one paper (passes 1 and 2 of three-pass reading)45 min
Saturday, 1 hourFinish pass 3 + write card notes1 hour
Sunday, 30 minFile notes, update your personal paper map30 min

Rhythm matters more than intensity. After 8 weeks, you'll have 8 card notes, a clear narrative thread, and an intuition for "which papers are worth reading."

8. 10 Questions to Ask Before Reading Any Paper ​

Use these 10 questions as a template before you open any paper. They'll quickly help you judge "whether to deep-dive and where to focus":

  1. What problem is it trying to solve? Is the problem real, or was it contrived just to justify the method?
  2. Why didn't previous baselines work? Without an answer to this, the method is likely self-indulgent.
  3. What's the core innovation? Can you state it in one sentence? If you can't, the authors haven't figured it out either.
  4. What assumptions does it make? Under what conditions do those assumptions hold? Data, compute, task distribution.
  5. How are experiments set up? Data, hyperparams, compute, evaluation protocol — miss one and you can't trust the conclusions.
  6. Does the improvement come from innovation or brute force? Check ablations: remove the novel component, see how much performance drops.
  7. What are the authors' stated limitations? Papers that thoughtfully write Limitations tend to be more credible.
  8. What experiments are missing? This is where improvement opportunities lie — and also interview bonus points.
  9. In my scenario, where is this most likely to fail? The core question for engineering deployment.
  10. What reusable things does it leave behind? Code, data, models, or just an intuition?

Why "pre-reading" questions?

You should write these down before opening the full text. Having your own questions before reading prevents you from passively following the authors' narrative — reading a paper is fundamentally a "conversation," not a "lecture."

9. Minimal Action Checklist ​

After reading this page, follow this checklist to get started in your first week:

text
□ Choose a path (A/B/C) and tell a friend your choice
□ Open [Reading Paths](/papers/paths) and note down your first paper
□ Open [Core Paper Deep Dives](/papers/core-papers) and read one section
□ Use "term-based navigation" to scan that paper's title, abstract, and figures
□ Complete the first pass of [three-pass reading](/papers/faq), write down three sentences
□ Set up your notes file (or Obsidian vault) and write your first card
□ Subscribe to arXiv's cs.CL / cs.LG (see [Following arXiv](/papers/faq))

10. The Most Important Piece of Advice ​

Take away mechanisms, not numbers

The numbers in papers are often the most memorable: "GPT-3 has 175B parameters," "GSM8K accuracy went from 18% to 57%." But remember:

Every absolute score depends on the data, compute, preprocessing, and evaluation protocol at the time. Change the data distribution, the hardware budget, or the evaluation script — and those numbers may no longer hold.

Three things are truly worth taking away:

  1. Mechanisms — Why does this method work? What structural problem does it solve? (e.g., FlashAttention solves the memory wall at the IO level for attention)
  2. Failure analysis — What experiments were missing? What approaches failed? What limitations did the authors state? Failures are often more valuable than successes.
  3. Boundary conditions — Under what conditions does this method hold? When does it fail? (e.g., CoT doesn't improve small models; it only emerges in sufficiently large models)

Leaderboard numbers go stale. Mechanisms and boundaries don't. This "read mechanisms, read boundaries" discipline is fully laid out in Reading Discipline & FAQ.

Further Reading ​

  • Reading Paths — The next step from here: choose your first path and get a specific paper checklist
  • What Is a Large Language Model — Domain concept foundation; we recommend building this global picture before reading papers
  • Paper Map — Want to build a macro coordinate system first? Start here
  • Glossary — Quick-term reference for paper reading
  • Evolution Timeline — A popular version of the paper timeline, cross-referenced with the map

References ​

All of these are real, publicly available resources for self-directed deepening: