Skip to content

Reading Discipline & FAQ

At a glance A complete methodology for reading LLM papers: the three-pass reading method with checklists for each pass, troubleshooting for when you can't understand, criteria for judging paper quality (including LLM-specific pitfalls), card note method, whether blogs alone suffice, tools and techniques for following arXiv, reproduction advice, and troubleshooting.

Reading Discipline & FAQ ​

One-sentence summary: This page is an operational manual for turning "reading papers" from a one-time burst of enthusiasm into a sustainable habit — three-pass reading, troubleshooting for when you get stuck, criteria for judging paper quality, card note method, and arXiv tracking workflows, all with actionable steps.

1. How to Read a Paper: The Three-Pass Method ​

This is the method from Karpathy's 2019 blog post combined with standard academic practice — three progressive layers, each with a clear deliverable:

text
Pass 1 (15 minutes): Grasp the skeleton — goal: "what is this paper about?"
├─ Read title → abstract → figures → conclusion
├─ Quickly skim figures and formula titles in the methods section (don't deep-read)
└─ Deliverable: Can answer "problem, method, effect" in three sentences

Pass 2 (1–2 hours): Dig into details — goal: "can I reproduce / apply it?"
├─ Deep-read methods + experiments + ablation sections
├─ Mark key formulas, hyperparameters, datasets, evaluation protocols
└─ Deliverable: Can answer "what's the core innovation, where it differs from baselines, what are the limitations"

Pass 3 (half day to several days): Pick it apart — goal: "can I improve it?"
├─ Reverse-derive formulas, identify the paper's hidden assumptions
├─ Design "what if..." controlled experiments
└─ Deliverable: Can state "improvement directions for this paper" and write them as notes

Checklist per Pass ​

PassCan Answer After ReadingWhat If You Can't?
Pass 1What's the problem? What's the method? What's the effect?Go back to abstract and figures; you must be able to answer three sentences within 10 minutes, otherwise return it to the reading queue
Pass 2Core innovation? How do baselines differ? What do ablations prove? What are the limitations?Cross-reference with the deep-dive of the same paper on Core Paper Deep Dives; see which layer you missed
Pass 3Hidden assumptions? Untested experiments? My improvement points?Discuss with a colleague who's read this paper, or write it down and reread it a few days later

When to stop

If after 15 minutes of Pass 1 you discover "this isn't relevant to my direction" — stop. Not every paper deserves a Pass 3. The judgment criteria are in Section 3.

2. What to Do When You Don't Understand ​

Not understanding is the norm, not a barrier. Treat symptoms:

SymptomReal CauseRemedy
Can't understand the abstractMissing domain backgroundRead the corresponding concept page first (e.g., Transformer Architecture, glossary), then come back
Can't push through formulasGaps in math foundationSkip formulas, read the text + read others' derivation notes; formulas aren't the main body of a paper
Can't read the figuresUnfamiliar with evaluation/dataRead the experimental setup text first; read figures alongside text
Can't grasp the methodMissing intuitionFind illustrated blogs, lecture videos (Karpathy/HF courses, see Awesome Resources)
Forget everything after readingNo output artifactsForce yourself to write a one-page card note (see Section 4); retention requires output
Too long to finishUnclear goalReturn to Reading Paths to set "how deep" for this paper; read only what you need to
Terminology overloadDidn't look up terms before readingScan the glossary first, or look up each unfamiliar term once — you'll be familiar after three encounters
Don't know why it mattersMissing coordinate systemCheck Paper Map to see which tributary this paper is in and who it inherits from

5-Minute Pre-Reading Prep ​

Not understanding often stems from no preparation. Spend 5 minutes before opening the full text:

text
□ Which institution/lab is this paper from? This determines credibility and resource level
□ Which tributary is it in on the [paper map](/papers/map)? What are the prerequisite concepts?
□ What problem can I guess from the title? Write it down
□ How many passes am I targeting? Set the goal
□ What prerequisite materials do I need (illustrated guides / videos / concept pages)? Open them first

The most wrong way to read

Starting from the first formula on page one and "reading linearly" for two hours, getting stuck in Section 3. A paper is not a novel — it's a reference manual — look at the map (abstract + figures) first, then decide how deep to go. This is the discipline emphasized from Getting Started.

3. How to Judge Paper Quality ​

A checklist for scoring papers (0/1 per item, higher score = more worth deep-reading):

DimensionGood Paper SignalsBad Paper Signals
ProblemReal, widespread, with a clear pain pointProblem is contrived just to justify the method
MethodMethod is concise, strongly tied to the problem, motivatedMethod is flashy, a patchwork of multiple tricks
ExperimentsFair baselines + complete ablations, code/data publicOnly compares SOTA numbers, no ablations, no code
NumbersReports confidence intervals / multiple runs, states eval protocolReports only the single best run, numbers aren't reproducible
HonestyThoughtfully writes Limitations, acknowledges failed experimentsEvades limitations, hides negative results
ImpactHeavily cited by subsequent work, paradigm-level contributionIgnored after release, or just hype

Five LLM-Specific Pitfalls ​

LLM papers have five categories of "domain-specific problems" to check separately:

PitfallPhenomenonHow to Spot It
Data contaminationTraining corpus leaked into test set (e.g., original GSM8K questions), inflating scoresCheck whether the paper discusses contamination; retest with fresh questions
Evaluation leakageUsing API models as judges, or benchmarks repeatedly grindedCheck evaluation protocol, judge setup and bias discussion
Compute opacityKey experiments used hidden scale / hidden budgetCheck compute descriptions; are ablations done at the same scale?
Prompt engineering maskingImprovement comes from carefully crafted prompts, not the method itselfCheck ablation for "standard prompt vs. paper prompt" comparison
Commercial-firstPaper trails behind product (e.g., some closed-source reports), incomplete evidenceCheck if only abstract-level evidence is given, lacking ablations

A special note

"Good" and "popular" are two different things for LLM papers. Some papers are shared wildly because their conclusions are explosive (e.g., "emergent abilities"), but their experimental evidence may be thin. Some papers are unremarkable but became cornerstones because they provided open-source models/data. When judging, first ask "what reusable thing did it leave for the community?"

4. How to Take Notes: Card Method ​

A long-lasting format — one card per paper, five sections. Five fields corresponding to "problem / method / numbers / limitations / my scenario" — both a post-reading summary and a future search index:

text
[Paper Card] Title: FlashAttention (Dao et al., 2022)
├─ Problem: Attention O(n²) memory, bottleneck is I/O not compute
├─ Method: tiling on SRAM + backprop recomputation, exact softmax
├─ Numbers: ~2–4× speedup (vs. standard PyTorch), memory O(n²)→O(n)
├─ Limitations: Only optimizes standard attention form, doesn't reduce compute complexity
└─ My scenario: Can long-context RAG retrieval benefit? → To be verified

Here's a second example for a more "conceptual" paper, showing both styles:

text
[Paper Card] Title: Chinchilla (Hoffmann et al., 2022)
├─ Problem: Scaling Laws only said "loss decreases with scale," not how to allocate params vs. data
├─ Method: Under fixed compute, search for optimal (N, D) pair; conclusion ~1:20 (20 tokens per param)
├─ Numbers: 70B/1.4T beat Gopher 280B/300B; better on average across 400+ tasks
├─ Limitations: Loss-optimal ≠ inference-cost-optimal; capability emergence doesn't directly correlate with loss
└─ My scenario: Use 1:20 to estimate data needs when choosing base models → Already written into selection checklist

Three note-taking rules:

  1. Write in your own words — no copying the abstract. If you can't, you didn't understand.
  2. Each card answers exactly five questions — no more, no fewer (five-section format above).
  3. Tag your cards (#architecture #alignment #systems…) — after accumulating 30, review by tag and you'll see your own understanding map.

How to use cards

Cards aren't collectibles. Spend 10 minutes weekly on "card review": pick an old card, try to retell it to yourself. If you can't, that paper needs a Pass 2 reread. What you can retell is yours; what you've written is just a candidate.

Tools are up to you: Obsidian, Notion, plain Markdown directories — the tool doesn't matter; the output structure does.

5. Reading Log Template: 10-Minute Weekly Dashboard ​

Cards are "per paper" records; logs are "weekly input/output" records. Keeping logs prevents "reading for nothing" — because the log forces you to answer: what did I invest, what did I produce, what questions remain?

text
[Weekly Log] 2025-W26 (Jun 23–29)
├─ This week's deep-dive: Chinchilla (Pass 3), RAG (Pass 2)
├─ This week's abstract reading: 14 arXiv papers (2 entered reading queue)
├─ Cards produced: 2 (#scale #retrieval)
├─ New questions:
│  ① Does Chinchilla's 1:20 still hold in the synthetic-data era?
│  ② Is there a systematic experiment on RAG chunk size vs. faithfulness?
├─ Next week plan: Deep-dive DPO (Pass 2) + run a minimal retrieval experiment
└─ One-line summary: This week's biggest cognitive update was...

Three usage rules:

  1. Write at a fixed time: 10 minutes every Sunday. Habit-building beats content perfection.
  2. Questions are the most valuable part: Questions accumulate into your next project/research direction — don't let them dissipate.
  3. Review the "one-line summary" column monthly: You'll see your cognition evolving — this is the biggest positive reinforcement for persistence.

6. Is Reading Only Blogs Enough? ​

Depends on the stage:

  • Beginner phase: Yes. Blogs, illustrated guides, and course videos are great "pre-reading" material — they lower your first barrier.
  • Practitioner phase: No. Blogs are secondhand: they can be stale, distorted, or lack failure details. To judge whether a method is production-ready, you must return to the original to check ablations, limitations, and evaluation protocols.
  • A compromise: Use blogs to build intuition, use papers to verify judgment — blogs first, originals second, not "either/or."

Blog vs. paper boundaries:

QuestionBlog handlesMust read paper
"What is this method roughly about?"Suitable—
"How much does this improve, and why?"—Suitable (check ablations and eval protocol)
"Would it work on my data?"—Suitable (check limitations and assumptions)
"What's new recently?"Suitable (as a lead)Follow leads with original papers

The reverse is also true

"Only reading papers and not blogs" is equally inefficient. Papers don't teach engineering intuition (how to tune hyperparams, how to avoid pitfalls) — that comes from practice guides and community blogs. Both hands must be strong.

7. How to Follow arXiv ​

arXiv publishes hundreds of new papers daily — browsing blindly will drown you. Here's an actionable filtering workflow:

text
Subscribe (5 min/day)
├─ Track categories: cs.CL (computational linguistics) + cs.LG (machine learning)
├─ Track keywords: Large Language Model / LLM / RLHF / Reasoning
├─ Filter standard: title + abstract only; 1 min to judge whether to add to "reading queue"
└─ Tools: arxiv-sanity, Hugging Face Papers, Papers with Code daily feed

Follow authors (10 min/2 weeks)
├─ Watch 4–6 labs/authors (OpenAI, DeepMind, Anthropic, Meta AI, target universities)
├─ Use X/GitHub/Google Scholar "follow author" for updates
└─ See if a new paper is "milestone-level" before deciding to deep-dive

Archive weekly (30 min)
├─ Weekly reading queue → triage by [Reading Paths](/papers/paths) three-path levels
├─ Deep-dive 1–2 papers, rest go into "abstracts read" list
└─ Update your own [paper map](/papers/map)

Tool and Technique Quick Reference ​

ToolPurposeTip
arXivPaper sourceSubscribe via cs.CL/cs.LG categories, or use RSS
arxiv-sanity / arxiv-sanity-litePaper ranking and similarity recommendationStar papers you follow — it learns your taste
Hugging Face PapersDaily curationSpend 2 min/day scanning headlines, add to "reading queue"
Papers with CodeReproduction status and codeFastest entry for judging "reproducibility"
Google ScholarCitation trackingWatch for citation count explosions — fastest way to spot hot topics
X / GitHubAuthor activityFollowing authors directly is the primary channel for frontier info

Following Citation Chains From One Paper ​

Once you find a good paper, use two citation chains to turn it into a batch of papers:

text
Go backwards (what it cites) → find the intellectual origin: read its Related Work, trace back to the foundational paper
Go forwards (who cited it) → find subsequent progress: check "Cited by" on Google Scholar / Semantic Scholar
Go sideways (same group, same topic) → find the full narrative: a series of work from the same team in the same direction

Practical advice: for every deep-dive paper, do at least one "go backwards" (understand who it stands on). For research, add "go forwards." This is an order of magnitude more efficient than aimlessly scrolling arXiv.

Don't fall into "update anxiety"

arXiv is an assembly line, not a bookshelf: the vast majority of new papers aren't worth deep-reading. Filter with "Section 3 criteria" — one to two deep-dives per week is already a professional researcher's pace. For the rest, abstracts are enough.

8. Reproduction Advice ​

Reproduction is the only way to verify "you really understand," but costs vary wildly:

Reproduction LevelContentCostRecommendation
Concept reproductionSmall-scale toy implementation (e.g., use nanoGPT to reproduce "small model loss decrease")HoursHighly recommended — must-do for beginners
Method reproductionReproduce a minimal version of LoRA / FlashAttentionDaysMust-do for advanced learners
Full reproductionReproduce all experiments in a paperWeeks to months + computeOnly for research
Massive-scale reproductionReproduce 70B / 1.4T trainingNot feasibleNot realistic — just read the tech report

Common Reproduction Failures and Troubleshooting ​

Failure SymptomMost Likely CauseTroubleshooting Direction
Loss doesn't decreaseLearning rate / data format errorFirst run the official minimal example, then swap in your data
Performance much worse than paperDataset / hyperparams / eval protocol mismatchCheck the paper's "experimental setup" checklist item by item
GPU memory OOMBatch size or sequence length exceededReduce batch/length, or go LoRA / quantization route
Unstable resultsRandom seeds, sampling params not fixedFix seeds, run multiple times, report mean and variance
Code doesn't match paperOfficial code has been iteratedTrust the paper; treat code as a reference implementation

Practical path: start with the nanoGPT handson from Building an LLM from Scratch (can run in a few hours), then reproduce a method paper (recommend LoRA, see Fine-Tuning Practice). One "read + write" loop beats reading ten papers.

9. FAQ Quick Reference ​

Here are the 10 most common questions about reading papers, one-page quick reference:

#QuestionOne-Sentence AnswerSee
1What if I only have 15 minutes a day?Do only Pass 1 (grasp skeleton); maintain trend sensitivitySection 6
2Must I understand all the math derivations?No; grasp mechanism and intuition first; derivations are Pass 3 materialSection 2
3How do I judge which papers are important?Check citation count + open-source status + team and experiment credibilitySection 3
4Do I need to subscribe to journals/conferences?No — arXiv + conference official proceedings is sufficientSection 7
5English reading is slow. What should I do?Read through the glossary first; reread classic papersSection 2
6How many papers to reach "beginner"?Deep-read 30 papers and write 30 cards to reach "can explain in depth"Section 4
7Can videos/lectures substitute for papers?Videos build intuition, papers build evidence — complementary, not substitutableSection 5
8Should I read survey papers?Yes; read surveys first to build the domain map, then read individual papersPaper Map
9How do I verify I really understand?Feynman technique: explain the five-section card clearly to someone else, without notesSection 4
10I always forget after reading. What should I do?Spaced repetition + card review. Output artifacts are memory anchorsSection 4

10. One Discipline ​

Reading discipline: the master rule

Read papers to build judgment, not to accumulate quantity. Ask yourself three questions weekly:

  1. How many papers did I truly understand this week (not just how many I read)?
  2. Can I explain its mechanism and boundaries in my own words?
  3. What verifiable impact did it have on my work or research?

If the answer to all three is "no," you didn't read deeply enough — go back to three-pass reading, reread, don't rush to the next paper.

One last psychological note

The frustration of reading papers (not understanding, forgetting, being left behind by the frontier) is systemic — it's not a personal failure. The three-piece set of methodology (this page) + rhythm (fixed weekly time) + output artifacts (cards) significantly alleviates it. People who've read 100 papers don't understand every one either — they're just better at deciding "which ones are worth understanding."

Further Reading ​

  • Getting Started — Section entry point: the three layers of value in paper reading
  • Reading Paths — After reading the methodology, return to define your path and checklist
  • Core Paper Deep Dives — The first batch of targets to apply the methods from this page
  • Paper Map — Build a coordinate system so your card notes have a place to land
  • Awesome Resources — Illustrated guides, courses, reproduction repos, and more supplementary ammo

References ​

Real, publicly available resources relevant to the methodology: