Skip to content

Reading Discipline & FAQ

Quick overview Methodology FAQ for reading papers: how to choose your first paper, what to do when you can't understand, the "three questions" note-taking framework, how to judge paper quality (vs. baseline, ablation, open-source code), how to reproduce, and how to present a paper well in 5 steps, plus common pitfalls and reading discipline. Organized in Q&A format for on-demand lookup.

Reading Discipline & FAQ ​

One-sentence definition: Reading discipline is a methodology for "how to read papers well" -- pick the right articles, push through the "can't understand" wall, leave mechanisms in notes, validate understanding through reproduction, and solidify expression through presentation. This page is organized in Q&A format -- look up what you need.

I. Reading Discipline: Set Rules First ​

The biggest enemy of reading papers isn't difficulty -- it's disorder. Four disciplines to start with:

  1. Position before close-reading: before reading a paper, ask "which main thread is it on in the Paper Map?" -- with coordinates, details have meaning;
  2. One paper at a time: grasping more than chewing leads to indigestion. 45 minutes per paper x multiple rounds beats sitting all afternoon chugging through;
  3. Read with questions: before each reading session, write down the questions you want answered (see the three questions in Section III), and only look for answers while reading;
  4. Always produce something after reading: if a paper doesn't produce at least one of notes / retelling / code, you haven't really read it.

A quick-reference table for daily use (detailed operations in corresponding sections):

ScenarioDisciplineSection
Before readingPosition first (Paper Map), then close-readSection I
Can't understandThree-pass layered reading, attack blockers next dayQ2
Taking notesThree questions: problem / contribution / ablationQ3
Judging qualityBaseline / ablation / open-source / data availableQ4
Reproducing a paperRead code first, fix seeds, record deviationsQ5
Presenting a paperPosition -> motivation -> method -> experiment -> critiqueQ6

II. Q1: How to Choose Your First Paper? ​

Work backward from the technology you're using. If you're writing PyTorch models, pick the most relevant one from Classic Papers Deep Dive (AlexNet or ResNet if using CNN, Attention Is All You Need if using Transformer). A paper suitable for a beginner looks like this:

  • Short (body <= 15 pages) with many diagrams;
  • Reproducible (has public datasets like MNIST/ImageNet/GLUE);
  • Clear mechanism (you can articulate "why it works," not just "it topped a leaderboard").

Avoid starting with engineering behemoths like AlphaFold2. Full reading order is at Reading Paths.

III. Q2: What to Do When You Can't Understand? ​

Use the "three-pass method" to read the paper in layers, instead of trying to understand it all at once:

  1. First pass (10 min): read only the title, abstract, conclusions, and figure captions -- answer "what did it do, what's the conclusion." Not understanding is normal; move straight to the second pass;
  2. Second pass (30-60 min): read the methods section, skip formulas you don't understand, and grasp the "input -> processing -> output" skeleton. Mark blocked points in red;
  3. Third pass (next day or later): come back to the blockers with your notes. Reading again the next day is a real technique -- after the brain offline-processes, many "dead knots" untangle themselves.

Handle specific blockers like this:

  • Don't understand a term: check the glossary first, then return to the original;
  • Can't follow a formula: fill in prerequisites with Math Primer (matrices, gradients, probability), derive only the formula's "skeleton," skip the details;
  • Unclear motivation: check the Paper Map to see who its "predecessors" were -- papers are answers to predecessors' pain points;
  • Still don't get it: read someone else's close reading or course notes, then return to the original. Second-hand interpretation is "preview," not "replacement."

Mindset on "not understanding"

Not understanding isn't a capability problem -- it's unsorted information hierarchy. Each paper is a compressed bar of 3 years of research; asking yourself to devour it in 2 hours will naturally make you choke. The three-pass method is about "swallow first -> digest later -> savor last."

IV. Q3: How to Take Notes? (Three-Questions Framework) ​

For each paper's notes, uniformly answer three questions, plus one more:

  1. What problem does it solve? -- background pain point, one sentence;
  2. What is the core contribution? -- the skeleton of method and mechanism, formulas allowed but with a "human language" version;
  3. What do the ablation experiments prove? -- what alternative explanations did the author rule out? which part truly drove results?
  4. (Optional) If I were doing this, what's next? -- write down one of your hypotheses.

A copyable template:

markdown
# Notes on "Title"
- Year / Authors / Venue: ___
- What problem does it solve: ___
- Core method (plain language): ___
- Key numbers: train/test set = ___, metric = ___ (remember "why it's credible," not "highest score")
- Ablation conclusions: ___
- My hypothesis / extension: ___
- To be filled: unfamiliar terms / formulas ___

This template is isomorphic with the four-part structure of Classic Papers Deep Dive -- do one paper with the template, and it gets faster from there.

V. Q4: How to Judge Whether a Paper Is Good? ​

Use four actionable criteria, reject "gut feeling":

  1. Compared to whom, and by how much? If there's no baseline comparison, or only claims "we're better" without numbers, be skeptical. Good papers compare multiple strong baselines and report both the magnitude of improvement and statistical significance;
  2. Were ablation experiments done? If a paper claims "our component A is critical" but never ablates A to see what happens to the loss, don't trust it. Ablation is the only evidence for judging "which designs truly matter" (methodology at Deep Learning Evaluation and Experiments);
  3. Is the code open-sourced? Open-source (weights/code) means reproducibility and testability -- this is the minimum bar of scientific integrity. "Breakthroughs" that only appear in press releases should default to discounted;
  4. Are the dataset and hyperparameters public? If a paper doesn't disclose training data scale, hyperparameters, or compute budget, you can't reproduce it and shouldn't -- question its value.

A complete approach to evaluating "does a technique actually work" (including evaluation set design and error analysis) is at Evaluation Practice.

VI. Q5: How to Reproduce a Paper? ​

Reproduction is the litmus test of understanding -- it precisely exposes every place where you thought you understood but didn't:

  1. Read code before writing code: find the author's open-source implementation, get it running first (check README, environment dependencies, data preparation). You're reproducing "conclusions," not "code";
  2. Start with the simplest sub-experiment: e.g., to reproduce ResNet, first get 20 layers running on CIFAR-10, then scale to ImageNet or deeper networks;
  3. Fix hyperparameters and seeds: the scores reported in a paper depend on the full recipe -- miss one trick and scores drop. This is the truth of "paper vs. engineering";
  4. Record every deviation: environment versions, batch sizes, learning rate schedulers -- all cause drift. Record them to know where the difference lies;
  5. When training goes wrong: first check the loss curve shape, then troubleshoot following the debugging and diagnostics flow, and cross-check recipe details with Training Recipes and Hyperparameter Tuning.

Reproduction Completion Criteria

Getting the metric within a reasonable range of the paper's reported value (+/- 1-2%) counts as success. Exact match is almost impossible -- papers omit more details than you think. The complete zero-to-reproduction workflow is at Build a Deep Learning Project from Scratch.

VII. Q6: How to Present a Paper? (5 Steps for Interview Talking Points) ​

Whether in an interview or a group meeting, use these 5 steps to present a paper in 5 minutes:

  1. One-sentence positioning: "This paper solves a YY problem in the XX domain, using method ZZ." -- 10 seconds to grab the audience;
  2. Background and motivation: where were predecessors, what's the pain point? (2-3 sentences, don't exceed 1 minute);
  3. Core method: explain the mechanism using "input -> key design -> output," ideally with a hand-drawn diagram (2 minutes);
  4. Key experiment: pick 1-2 most persuasive experiments, report numbers but emphasize conclusions ("X points higher than the strongest baseline on XX, ablation shows component A contributed the most");
  5. Limitations and your critique: what's imperfect about this paper? what would you do differently? -- this step is the watershed between "read it" and "understood it."

Practice materials at Interview Question Bank; outputs can be included in Portfolio Projects as project experience.

VIII. Q7: What Are the Common Pitfalls? ​

  • Drawing conclusions from the abstract alone: abstracts are written by authors, so they're naturally self-congratulatory. Conclusions should be based on experiments and methods;
  • Only reading Chinese summaries, never the original: second-hand info loses critical details or even transmits wrong conclusions. Summaries are for preview; the original is the main course;
  • Memorizing numbers, not mechanisms: error rates and parameter counts expire, mechanisms don't (see the "take away mechanisms" advice on the home page);
  • Greed over depth: "skimming" 20 papers in a week is worse than close-reading 2. Depth beats quantity;
  • Skipping "limitations and subsequent impact": this section often holds the key to understanding a paper's true value, and is the part that shines most in interviews;
  • Not reproducing: reading 100 papers is worth less than running one end-to-end. One reproduction beats ten browse-readings;
  • Treating paper conclusions as eternal truths: every conclusion has boundary conditions (datasets, hyperparameters, compute). Before transferring to a new scenario, first question its scope of applicability. Related anti-patterns at Common Pitfalls and Anti-Patterns.

IX. Boundaries and Trade-offs ​

  • Applicability of methodology: the three-pass method and three-question notes work for "algorithm/architecture" papers; pure theory papers (proof-heavy) and engineering reports (system-heavy) need tweaks -- the former needs more math, the latter needs more system design reading.
  • Discipline vs. flexibility: rules are your starting point, not your ending point. When you're "in the flow," don't mechanically follow steps -- discipline serves flow state, not the other way around.
  • Return on investment: reading papers is high-investment, high-return, but for learners who just want to "use tools," first run through Progressive Tutorial: Three Versions, Get It Running before entering papers -- more cost-effective than starting from papers.
  • Time-sensitive content: for discussions involving the latest models or leaderboards, remember the timeliness notice in Frontier Advances, and mark the cutoff date when citing.

Further Reading ​

References ​