Theme
Reading Discipline & FAQ
One-sentence summary: This page is an operational manual for turning "reading papers" from a one-time burst of enthusiasm into a sustainable habit — three-pass reading, troubleshooting for when you get stuck, criteria for judging paper quality, card note method, and arXiv tracking workflows, all with actionable steps.
1. How to Read a Paper: The Three-Pass Method
This is the method from Karpathy's 2019 blog post combined with standard academic practice — three progressive layers, each with a clear deliverable:
text
Pass 1 (15 minutes): Grasp the skeleton — goal: "what is this paper about?"
├─ Read title → abstract → figures → conclusion
├─ Quickly skim figures and formula titles in the methods section (don't deep-read)
└─ Deliverable: Can answer "problem, method, effect" in three sentences
Pass 2 (1–2 hours): Dig into details — goal: "can I reproduce / apply it?"
├─ Deep-read methods + experiments + ablation sections
├─ Mark key formulas, hyperparameters, datasets, evaluation protocols
└─ Deliverable: Can answer "what's the core innovation, where it differs from baselines, what are the limitations"
Pass 3 (half day to several days): Pick it apart — goal: "can I improve it?"
├─ Reverse-derive formulas, identify the paper's hidden assumptions
├─ Design "what if..." controlled experiments
└─ Deliverable: Can state "improvement directions for this paper" and write them as notesChecklist per Pass
| Pass | Can Answer After Reading | What If You Can't? |
|---|---|---|
| Pass 1 | What's the problem? What's the method? What's the effect? | Go back to abstract and figures; you must be able to answer three sentences within 10 minutes, otherwise return it to the reading queue |
| Pass 2 | Core innovation? How do baselines differ? What do ablations prove? What are the limitations? | Cross-reference with the deep-dive of the same paper on Core Paper Deep Dives; see which layer you missed |
| Pass 3 | Hidden assumptions? Untested experiments? My improvement points? | Discuss with a colleague who's read this paper, or write it down and reread it a few days later |
When to stop
If after 15 minutes of Pass 1 you discover "this isn't relevant to my direction" — stop. Not every paper deserves a Pass 3. The judgment criteria are in Section 3.
2. What to Do When You Don't Understand
Not understanding is the norm, not a barrier. Treat symptoms:
| Symptom | Real Cause | Remedy |
|---|---|---|
| Can't understand the abstract | Missing domain background | Read the corresponding concept page first (e.g., Transformer Architecture, glossary), then come back |
| Can't push through formulas | Gaps in math foundation | Skip formulas, read the text + read others' derivation notes; formulas aren't the main body of a paper |
| Can't read the figures | Unfamiliar with evaluation/data | Read the experimental setup text first; read figures alongside text |
| Can't grasp the method | Missing intuition | Find illustrated blogs, lecture videos (Karpathy/HF courses, see Awesome Resources) |
| Forget everything after reading | No output artifacts | Force yourself to write a one-page card note (see Section 4); retention requires output |
| Too long to finish | Unclear goal | Return to Reading Paths to set "how deep" for this paper; read only what you need to |
| Terminology overload | Didn't look up terms before reading | Scan the glossary first, or look up each unfamiliar term once — you'll be familiar after three encounters |
| Don't know why it matters | Missing coordinate system | Check Paper Map to see which tributary this paper is in and who it inherits from |
5-Minute Pre-Reading Prep
Not understanding often stems from no preparation. Spend 5 minutes before opening the full text:
text
□ Which institution/lab is this paper from? This determines credibility and resource level
□ Which tributary is it in on the [paper map](/papers/map)? What are the prerequisite concepts?
□ What problem can I guess from the title? Write it down
□ How many passes am I targeting? Set the goal
□ What prerequisite materials do I need (illustrated guides / videos / concept pages)? Open them firstThe most wrong way to read
Starting from the first formula on page one and "reading linearly" for two hours, getting stuck in Section 3. A paper is not a novel — it's a reference manual — look at the map (abstract + figures) first, then decide how deep to go. This is the discipline emphasized from Getting Started.
3. How to Judge Paper Quality
A checklist for scoring papers (0/1 per item, higher score = more worth deep-reading):
| Dimension | Good Paper Signals | Bad Paper Signals |
|---|---|---|
| Problem | Real, widespread, with a clear pain point | Problem is contrived just to justify the method |
| Method | Method is concise, strongly tied to the problem, motivated | Method is flashy, a patchwork of multiple tricks |
| Experiments | Fair baselines + complete ablations, code/data public | Only compares SOTA numbers, no ablations, no code |
| Numbers | Reports confidence intervals / multiple runs, states eval protocol | Reports only the single best run, numbers aren't reproducible |
| Honesty | Thoughtfully writes Limitations, acknowledges failed experiments | Evades limitations, hides negative results |
| Impact | Heavily cited by subsequent work, paradigm-level contribution | Ignored after release, or just hype |
Five LLM-Specific Pitfalls
LLM papers have five categories of "domain-specific problems" to check separately:
| Pitfall | Phenomenon | How to Spot It |
|---|---|---|
| Data contamination | Training corpus leaked into test set (e.g., original GSM8K questions), inflating scores | Check whether the paper discusses contamination; retest with fresh questions |
| Evaluation leakage | Using API models as judges, or benchmarks repeatedly grinded | Check evaluation protocol, judge setup and bias discussion |
| Compute opacity | Key experiments used hidden scale / hidden budget | Check compute descriptions; are ablations done at the same scale? |
| Prompt engineering masking | Improvement comes from carefully crafted prompts, not the method itself | Check ablation for "standard prompt vs. paper prompt" comparison |
| Commercial-first | Paper trails behind product (e.g., some closed-source reports), incomplete evidence | Check if only abstract-level evidence is given, lacking ablations |
A special note
"Good" and "popular" are two different things for LLM papers. Some papers are shared wildly because their conclusions are explosive (e.g., "emergent abilities"), but their experimental evidence may be thin. Some papers are unremarkable but became cornerstones because they provided open-source models/data. When judging, first ask "what reusable thing did it leave for the community?"
4. How to Take Notes: Card Method
A long-lasting format — one card per paper, five sections. Five fields corresponding to "problem / method / numbers / limitations / my scenario" — both a post-reading summary and a future search index:
text
[Paper Card] Title: FlashAttention (Dao et al., 2022)
├─ Problem: Attention O(n²) memory, bottleneck is I/O not compute
├─ Method: tiling on SRAM + backprop recomputation, exact softmax
├─ Numbers: ~2–4× speedup (vs. standard PyTorch), memory O(n²)→O(n)
├─ Limitations: Only optimizes standard attention form, doesn't reduce compute complexity
└─ My scenario: Can long-context RAG retrieval benefit? → To be verifiedHere's a second example for a more "conceptual" paper, showing both styles:
text
[Paper Card] Title: Chinchilla (Hoffmann et al., 2022)
├─ Problem: Scaling Laws only said "loss decreases with scale," not how to allocate params vs. data
├─ Method: Under fixed compute, search for optimal (N, D) pair; conclusion ~1:20 (20 tokens per param)
├─ Numbers: 70B/1.4T beat Gopher 280B/300B; better on average across 400+ tasks
├─ Limitations: Loss-optimal ≠ inference-cost-optimal; capability emergence doesn't directly correlate with loss
└─ My scenario: Use 1:20 to estimate data needs when choosing base models → Already written into selection checklistThree note-taking rules:
- Write in your own words — no copying the abstract. If you can't, you didn't understand.
- Each card answers exactly five questions — no more, no fewer (five-section format above).
- Tag your cards (#architecture #alignment #systems…) — after accumulating 30, review by tag and you'll see your own understanding map.
How to use cards
Cards aren't collectibles. Spend 10 minutes weekly on "card review": pick an old card, try to retell it to yourself. If you can't, that paper needs a Pass 2 reread. What you can retell is yours; what you've written is just a candidate.
Tools are up to you: Obsidian, Notion, plain Markdown directories — the tool doesn't matter; the output structure does.
5. Reading Log Template: 10-Minute Weekly Dashboard
Cards are "per paper" records; logs are "weekly input/output" records. Keeping logs prevents "reading for nothing" — because the log forces you to answer: what did I invest, what did I produce, what questions remain?
text
[Weekly Log] 2025-W26 (Jun 23–29)
├─ This week's deep-dive: Chinchilla (Pass 3), RAG (Pass 2)
├─ This week's abstract reading: 14 arXiv papers (2 entered reading queue)
├─ Cards produced: 2 (#scale #retrieval)
├─ New questions:
│ ① Does Chinchilla's 1:20 still hold in the synthetic-data era?
│ ② Is there a systematic experiment on RAG chunk size vs. faithfulness?
├─ Next week plan: Deep-dive DPO (Pass 2) + run a minimal retrieval experiment
└─ One-line summary: This week's biggest cognitive update was...Three usage rules:
- Write at a fixed time: 10 minutes every Sunday. Habit-building beats content perfection.
- Questions are the most valuable part: Questions accumulate into your next project/research direction — don't let them dissipate.
- Review the "one-line summary" column monthly: You'll see your cognition evolving — this is the biggest positive reinforcement for persistence.
6. Is Reading Only Blogs Enough?
Depends on the stage:
- Beginner phase: Yes. Blogs, illustrated guides, and course videos are great "pre-reading" material — they lower your first barrier.
- Practitioner phase: No. Blogs are secondhand: they can be stale, distorted, or lack failure details. To judge whether a method is production-ready, you must return to the original to check ablations, limitations, and evaluation protocols.
- A compromise: Use blogs to build intuition, use papers to verify judgment — blogs first, originals second, not "either/or."
Blog vs. paper boundaries:
| Question | Blog handles | Must read paper |
|---|---|---|
| "What is this method roughly about?" | Suitable | — |
| "How much does this improve, and why?" | — | Suitable (check ablations and eval protocol) |
| "Would it work on my data?" | — | Suitable (check limitations and assumptions) |
| "What's new recently?" | Suitable (as a lead) | Follow leads with original papers |
The reverse is also true
"Only reading papers and not blogs" is equally inefficient. Papers don't teach engineering intuition (how to tune hyperparams, how to avoid pitfalls) — that comes from practice guides and community blogs. Both hands must be strong.
7. How to Follow arXiv
arXiv publishes hundreds of new papers daily — browsing blindly will drown you. Here's an actionable filtering workflow:
text
Subscribe (5 min/day)
├─ Track categories: cs.CL (computational linguistics) + cs.LG (machine learning)
├─ Track keywords: Large Language Model / LLM / RLHF / Reasoning
├─ Filter standard: title + abstract only; 1 min to judge whether to add to "reading queue"
└─ Tools: arxiv-sanity, Hugging Face Papers, Papers with Code daily feed
Follow authors (10 min/2 weeks)
├─ Watch 4–6 labs/authors (OpenAI, DeepMind, Anthropic, Meta AI, target universities)
├─ Use X/GitHub/Google Scholar "follow author" for updates
└─ See if a new paper is "milestone-level" before deciding to deep-dive
Archive weekly (30 min)
├─ Weekly reading queue → triage by [Reading Paths](/papers/paths) three-path levels
├─ Deep-dive 1–2 papers, rest go into "abstracts read" list
└─ Update your own [paper map](/papers/map)Tool and Technique Quick Reference
| Tool | Purpose | Tip |
|---|---|---|
| arXiv | Paper source | Subscribe via cs.CL/cs.LG categories, or use RSS |
| arxiv-sanity / arxiv-sanity-lite | Paper ranking and similarity recommendation | Star papers you follow — it learns your taste |
| Hugging Face Papers | Daily curation | Spend 2 min/day scanning headlines, add to "reading queue" |
| Papers with Code | Reproduction status and code | Fastest entry for judging "reproducibility" |
| Google Scholar | Citation tracking | Watch for citation count explosions — fastest way to spot hot topics |
| X / GitHub | Author activity | Following authors directly is the primary channel for frontier info |
Following Citation Chains From One Paper
Once you find a good paper, use two citation chains to turn it into a batch of papers:
text
Go backwards (what it cites) → find the intellectual origin: read its Related Work, trace back to the foundational paper
Go forwards (who cited it) → find subsequent progress: check "Cited by" on Google Scholar / Semantic Scholar
Go sideways (same group, same topic) → find the full narrative: a series of work from the same team in the same directionPractical advice: for every deep-dive paper, do at least one "go backwards" (understand who it stands on). For research, add "go forwards." This is an order of magnitude more efficient than aimlessly scrolling arXiv.
Don't fall into "update anxiety"
arXiv is an assembly line, not a bookshelf: the vast majority of new papers aren't worth deep-reading. Filter with "Section 3 criteria" — one to two deep-dives per week is already a professional researcher's pace. For the rest, abstracts are enough.
8. Reproduction Advice
Reproduction is the only way to verify "you really understand," but costs vary wildly:
| Reproduction Level | Content | Cost | Recommendation |
|---|---|---|---|
| Concept reproduction | Small-scale toy implementation (e.g., use nanoGPT to reproduce "small model loss decrease") | Hours | Highly recommended — must-do for beginners |
| Method reproduction | Reproduce a minimal version of LoRA / FlashAttention | Days | Must-do for advanced learners |
| Full reproduction | Reproduce all experiments in a paper | Weeks to months + compute | Only for research |
| Massive-scale reproduction | Reproduce 70B / 1.4T training | Not feasible | Not realistic — just read the tech report |
Common Reproduction Failures and Troubleshooting
| Failure Symptom | Most Likely Cause | Troubleshooting Direction |
|---|---|---|
| Loss doesn't decrease | Learning rate / data format error | First run the official minimal example, then swap in your data |
| Performance much worse than paper | Dataset / hyperparams / eval protocol mismatch | Check the paper's "experimental setup" checklist item by item |
| GPU memory OOM | Batch size or sequence length exceeded | Reduce batch/length, or go LoRA / quantization route |
| Unstable results | Random seeds, sampling params not fixed | Fix seeds, run multiple times, report mean and variance |
| Code doesn't match paper | Official code has been iterated | Trust the paper; treat code as a reference implementation |
Practical path: start with the nanoGPT handson from Building an LLM from Scratch (can run in a few hours), then reproduce a method paper (recommend LoRA, see Fine-Tuning Practice). One "read + write" loop beats reading ten papers.
9. FAQ Quick Reference
Here are the 10 most common questions about reading papers, one-page quick reference:
| # | Question | One-Sentence Answer | See |
|---|---|---|---|
| 1 | What if I only have 15 minutes a day? | Do only Pass 1 (grasp skeleton); maintain trend sensitivity | Section 6 |
| 2 | Must I understand all the math derivations? | No; grasp mechanism and intuition first; derivations are Pass 3 material | Section 2 |
| 3 | How do I judge which papers are important? | Check citation count + open-source status + team and experiment credibility | Section 3 |
| 4 | Do I need to subscribe to journals/conferences? | No — arXiv + conference official proceedings is sufficient | Section 7 |
| 5 | English reading is slow. What should I do? | Read through the glossary first; reread classic papers | Section 2 |
| 6 | How many papers to reach "beginner"? | Deep-read 30 papers and write 30 cards to reach "can explain in depth" | Section 4 |
| 7 | Can videos/lectures substitute for papers? | Videos build intuition, papers build evidence — complementary, not substitutable | Section 5 |
| 8 | Should I read survey papers? | Yes; read surveys first to build the domain map, then read individual papers | Paper Map |
| 9 | How do I verify I really understand? | Feynman technique: explain the five-section card clearly to someone else, without notes | Section 4 |
| 10 | I always forget after reading. What should I do? | Spaced repetition + card review. Output artifacts are memory anchors | Section 4 |
10. One Discipline
Reading discipline: the master rule
Read papers to build judgment, not to accumulate quantity. Ask yourself three questions weekly:
- How many papers did I truly understand this week (not just how many I read)?
- Can I explain its mechanism and boundaries in my own words?
- What verifiable impact did it have on my work or research?
If the answer to all three is "no," you didn't read deeply enough — go back to three-pass reading, reread, don't rush to the next paper.
One last psychological note
The frustration of reading papers (not understanding, forgetting, being left behind by the frontier) is systemic — it's not a personal failure. The three-piece set of methodology (this page) + rhythm (fixed weekly time) + output artifacts (cards) significantly alleviates it. People who've read 100 papers don't understand every one either — they're just better at deciding "which ones are worth understanding."
Further Reading
- Getting Started — Section entry point: the three layers of value in paper reading
- Reading Paths — After reading the methodology, return to define your path and checklist
- Core Paper Deep Dives — The first batch of targets to apply the methods from this page
- Paper Map — Build a coordinate system so your card notes have a place to land
- Awesome Resources — Illustrated guides, courses, reproduction repos, and more supplementary ammo
References
Real, publicly available resources relevant to the methodology:
- Karpathy. A Recipe for Training Neural Networks (2019) — One of the original sources of the three-pass reading method; "the engineering attitude toward reading papers"
- Karpathy. Let's build GPT (2023) — Coding along while deep-reading the GPT paper; a model of reproduction-style learning
- arXiv — The primary paper publishing venue; pair with
cs.CL/cs.LGcategory subscriptions - arxiv-sanity (Karpathy) — arXiv paper filtering and ranking tool
- Hugging Face Papers — Daily curated paper leaderboard, great for quickly scanning headlines
- Papers with Code — Paper + code + reproduction status aggregation; the fastest entry for judging "reproducibility"