Skip to content

Reading Discipline & FAQ

Quick overview A methodology manual for reading papers: how to read numbers and figures in papers, how to use red-flag signals to judge paper quality, plus 12 FAQ answers—can't understand, can't remember, can't find related papers, can't finish them all: all have concrete countermeasures.

Reading Discipline & FAQ ​

In one sentence: What matters most in reading papers isn't "how many you've read" but "whether you have a repeatable reading discipline"—knowing which numbers to trust, which figures to see through, and which papers to throw away outright. This article establishes discipline first, then answers all your most frequent questions from "can't understand" to "can't finish."

How to use this article

Read the discipline section (Section 1) once as a deep read, then use the "red-flag signal table" to self-check each time you read a paper; the FAQ section doesn't need to be read cover to cover—come back to it when you hit a corresponding problem. Treat it as a methodology toolkit, not an article you must finish in one sitting.

1. Reading Discipline: Numbers, Figures, and the Paper's Referee ​

1. How to Read Numbers: Every Score Is a Conspiracy of "model + data + compute" ​

The most eye-catching thing in a paper is numbers: "top-1 accuracy of 88.5%," "BLEU 41.8," "beats SOTA by 2.3 points on X." But remember one iron rule:

A benchmark score is never an absolute value of "how good a method is"—it's the joint product of "method + data + compute + evaluation protocol." Change any one condition, and the numbers become incomparable.

So whenever you see any number, force yourself to ask four questions:

Follow-up QuestionSpecific MeaningWhy It Matters
DataWhich dataset? Which split? Any data augmentation?ImageNet's official split and community re-splits can differ by over 1 point
ComputeHow many GPUs? How long to train? How large was the batch size?The same method with twice the compute often brings notable gains
EvaluationHow is the metric calculated? Single run or averaged over multiple runs? Standard deviation reported?Only reporting the best run, without variance, may just be luck
ComparisonAre baselines implemented fairly? Were baselines tuned?"Strawman" baseline comparison is the most common paper trick

Pay special attention to relative improvement vs. absolute improvement: when a paper says "improved by 30% on X," it could be from 0.2% to 0.26% (essentially meaningless in absolute terms), or from 90% to 90.26% (already notable on a saturated task). Always convert to absolute numbers before making a judgment.

Another number trap: data contamination. In the era of large models, test sets may already appear in the model's pretraining data—a model that's "god-like" on a benchmark may simply have memorized the answers, not learned reasoning. The way to judge is to check whether the paper discloses the relationship between training data and test sets, and how the model performs on small-sample, out-of-distribution data.

One-sentence memory hook

Take mechanisms, not numbers. Numbers are products of "that time, that place, that machine"; mechanisms (why it works, where boundaries lie) are what transcend time and space. This is the same reading philosophy repeatedly emphasized in the Start Here section.

Figures are the highest information-density part of a paper and the most easily misled by "seeing what you want to see." Read in three steps:

Step 1: Look at axes    Is the x-axis log or linear? What's the y-axis range? What are the units?
                    (log coordinates "flatten" differences; linear coordinates amplify early jitter)

Step 2: Look at trends  Is the curve monotonically rising or U-shaped after overfitting?
                    Are methods "spreading apart" or "running neck and neck"?

Step 3: Look at conclusions  What conclusion does the author draw from this figure?
                    Does the figure actually support this conclusion? Could it be drawn more honestly?

Three frequent "figure reading" traps:

  • Only reporting improvement curves, not absolute values: In a plot of loss vs. training steps, if the y-axis origin is artificially truncated, the curve looks much steeper. Watch the axis range.
  • Ablation figures: look at "who hurts most when removed": The meaning of ablation experiments is to answer "how much does each component contribute." The reading approach isn't to see "how strong the full model is" but to see which component's removal causes the biggest drop—that's the method's true core innovation. Components with insignificant drops are often just nice-to-have extras, or even dispensable.
  • Scaling law figures: look at slope and saturation: In the deep learning era, double-log plots of "performance vs. parameters/data" are common. Look at whether the trend is still rising and whether saturation has appeared, not the specific value at a single point. A curve that can still climb is more informative than a "high score" that's already flat.

3. How to Judge Paper Quality: Red-Flag Signal Table ​

After reading enough papers, you'll notice that good and bad papers diverge clearly in "degree of honesty." Below is a table of signals worth watching—the more that appear, the more to lower your trust level:

Red-Flag SignalManifestationCountermeasure
Only reporting numbers, no protocolClaims "achieved SOTA" but doesn't disclose dataset splits, evaluation scripts, or number of runsCheck code/appendix directly; if unavailable, flag as "non-reproducible"
Strawman baseline comparisonOnly compares against "untuned old methods," avoiding the strongest contemporariesCheck Papers with Code for the real SOTA on that task
No ablation experimentsMethod has five or six modules but no "what if we remove each one" analysisCore contribution can't be pinpointed; innovation credibility is questionable
Single datasetOnly reports results on ImageNet (or one benchmark)Very likely overfitting to one benchmark
No variance reportedAll experiments report "the best run," no mean, no standard deviationGains from small random seeds aren't trustworthy
No discussion of limitationsNo Discussion/Limitation section anywhereEither the authors haven't thought it through, or they're hiding problems
Abstract uses relative numbers"Improved by 30%," "significantly surpasses" but no absolute numbers in the bodyLikely base is too small or comparison is unfair
Code always "coming soon"Paper says "Code will be released" but no repo is foundReproducibility is the lowest standard of science; be extra cautious without code

Red-flag signals ≠ the paper is bad

Red-flag signals are indicators of "raise your vigilance," not a "shoot on sight" verdict. A paper with weaknesses can still contribute valuable ideas—maybe its method doesn't hold under other settings, but its problem definition and approach are correct. The right posture: flag the red flags in your notes, assign a "credibility grade" to the paper after reading, rather than thinking in black and white.

2. FAQ: Twelve Questions on the Paper-Reading Journey ​

Q1. How much math do I need before reading papers? ​

Less than you think. You don't need to master all the math in a paper—what you need is basic competency in "understanding what symbols are doing": vector and matrix operations, calculus derivatives and the chain rule, probability expectations and conditional probability, plus a touch of information theory (entropy, KL divergence). All of this is covered in the Math Primer. The real skill is layered consumption of math: Pass 1 reads conclusions without derivations; Pass 2 grasps "what the formula computes, what the I/O is"; only when you plan to reproduce or improve the method do you need line-by-line derivations. Most beginner papers (classic models, architecture innovations) don't have high math density—what's high is concept density.

Q2. What if I can't understand? ​

First, accept a fact: not understanding is the default state, not an anomaly. Papers aren't textbooks; authors assume readers are peers and skip a lot of background. The correct handling order: ① swap materials—there are usually blog explanations, YouTube walkthroughs, and visual notes for any paper; use secondhand materials first to build a big picture (Jay Alammar's visual series, distill.pub are excellent crutches); ② go back and fill prerequisites—can't understand often isn't this paper's problem but missing prior knowledge; follow the citation chain to read foundational papers it references (can't understand ResNet? go back and read AlexNet; can't understand Transformer? go back and read the original attention paper); ③ mark, skip, and return—note what you don't understand, read what you can first, and many questions dissolve themselves on a second read. Reading one paper three times is far more effective than reading three papers once each.

Q3. Should I read the English original? ​

Yes, and as early as possible. Three reasons: ① translation causes information loss—terms, sentence structures, and the "hedges" unique to papers (the subtext behind "surprisingly" or "to the best of our knowledge") are only conveyed in the original; ② the ML field has no reliable complete Chinese translation ecosystem; frontier papers are always English-first, and Chinese explanations often lag by months or don't exist at all; ③ reading the original is a "one-time investment with long-term payoff"—after reading 20 English papers, your reading speed will have a qualitative leap because paper English is highly patterned. Advice for beginners: "read Chinese and English mixed"—first read Chinese surveys/blogs to build a big picture, then return to the English original to check details, treating the original as the only source of truth.

Q4. How to take notes? ​

The goal of note-taking is "still understandable in three months," not "feeling like I understood at the time." I recommend a structured "one-page note" template:

Paper title / year / venue
One-sentence: what problem does it solve? (in your own words)
Core mechanism: what's the core idea? (1–3 sentences + one diagram you draw yourself)
Key numbers: the 2–3 most important results + dataset/protocol
Ablation conclusions: which component contributes the most?
Limitations: self-admitted ones + ones you spotted
Relation to me: what's the use for my project/learning?
Credibility grade: high / medium / low + why

Strongly recommend hand-drawing mechanism diagrams—draw boxes and arrows to sketch the paper's method. If you can draw it, you truly understand it. Note-taking tools don't matter—Obsidian, Notion, plain Markdown all work—but the key is giving every paper a fixed template so notes are batch-searchable. Each entry in the site's Classic Paper Deep Dives is written in the fixed structure of "problem background → core mechanism → experimental conclusions → implications"—you can use them directly as your note template.

Use the "start from one paper, bidirectional spread" method—far more efficient than random search box queries:

Paper on Method X
   ├─ Forward: read its References ──▶ find earlier source papers
   └─ Backward: read who cited it ──▶ find follow-up improvement work

Specific tools: Google Scholar for citations and related articles; Semantic Scholar has good API and recommendations; Connected Papers visualizes the paper's citation network graphically, ideal for quickly locating "the core paper circle of a direction." Papers with Code is the best entry for "task → SOTA → code → paper"—follow a benchmark's leaderboard to find all major methods for that task. Combined with the site's Paper Map, first build a global coordinate system, then let citation chains take you deeper.

Q6. How to judge when a paper is outdated? ​

Three judgment dimensions: ① Citations and follow-up work—check its citation count on Semantic Scholar or Google Scholar and who's citing it; if follow-up work has clearly overturned or significantly improved on it, downgrade; ② Hardware and compute context—AlexNet in 2012 trained on two GTX 580s for days; base models in the 2020s deploy on tens of thousands of GPUs—the same method under different compute contexts can yield completely different conclusions. Watch whether the paper's "compute premise" still holds; ③ Surveys and evaluation benchmarks—if a task's latest survey no longer mentions this paper, or the latest benchmark's SOTA uses a completely different paradigm (e.g., from feature engineering to representation learning), it has entered "historical value" status. The meaning of judging outdatedness isn't "don't read old papers"—it's to give them the right positioning—old papers are often the keys to understanding new ones, which is exactly how Classic Paper Deep Dives and Frontier Progress are organized along this "past-to-present" approach.

Q7. Should I reproduce experiments? ​

It depends on your goal, three intensities:

GoalReproduction IntensityDescription
Beginner learningNo reproduction neededUnderstand + notes + explain to others is enough; spend time on breadth
Production deploymentMust reproduce (on small-scale data)Until you run it, you'll never know how many "hidden details" (learning rate scheduling, data augmentation, initialization) the paper didn't spell out
Doing researchDeep reproduction + ablationThis is the ticket to research: without reproduction, follow-up improvements are castles in the air

Key insight: the cost of reproducing one paper is often the true measure of that paper's value. If a method's "official code can't reproduce the paper numbers," that information itself is more valuable than the paper. In the 2020s, prioritize papers with official open-source implementations—use the awesome resource list or Papers with Code to find official code, saving an order of magnitude of time versus building from scratch.

Q8. Too many papers, can't finish them all? ​

First, break an illusion: "finishing all papers" was never the goal; not finishing doesn't mean failure. Papers are "on-demand" resources, not "must-clear-a-queue" checklists. Three practical rules: ① the 80/20 rule—20% of papers deserve deep reading, 80% just need "read abstract + look at figures + write one sentence"—deep reading takes a serious one-hour sit-down; skim takes ten minutes. ② read by task, not by heat—give yourself a current theme (e.g., "attention variants of Transformers"), only read papers in that theme, write a summary before switching. Far more effective than "randomly brushing two papers today." ③ tolerate "read and forget"—the brain isn't a database; the key is making notes your external memory. For specific route planning, see Reading Paths; don't be greedy, no more than two or three papers deep-read simultaneously.

Q9. How to explain a paper to others? ​

The Feynman technique is best practice for paper reading. Before explaining, ask yourself: can I explain in three sentences "what problem this paper solves, how it solves it, and how well it works?" If not, you don't understand yet. Specific operations:

  1. Explain to a 5-year-old (or your non-technical friend): only explain motivation and intuition, no jargon—forces you to find the method's core intuition in the plainest language.
  2. Explain to peers (colleagues/classmates/interviewers): explain mechanism, experiments, limitations—forces you to think through every design choice and its alternatives.
  3. Write a 300-word summary: explaining is easier to fudge than writing. Getting it written reveals where you're blocked.

The interview scenario is particularly useful: "discussing a paper thoroughly" is virtually guaranteed in algorithm role interviews, with the standard answer structure being "problem → intuition → mechanism → experiments → limitations → what would I do." If you can't explain it, that's because the "relation to me" section of your notes is empty.

Q10. How to quickly judge whether a paper is worth reading? ​

Before deep reading, first filter with a "three-minute funnel":

Step 1  Title + abstract (30 sec)  →  Does the problem relate to me? Can the method be explained in one sentence?
Step 2  Figures + conclusion (60 sec)  →  Is the improvement significant? Under what protocol?
Step 3  Last paragraph of introduction (30 sec)  →  Are the author's contribution list specific?
Step 4  Check code/reproduction status (60 sec)  →  Is there official code? How many stars? Has anyone reproduced it successfully?

If two or more of these questions have no answer or a negative answer, file it under "to-read." Judging "not worth reading" is as important as judging "worth reading"—save limited time for truly important papers. Don't be hijacked by the word "SOTA"—leaderboards change monthly, and a paper's value extends far beyond its ranking.

Q11. Read surveys or read original papers? ​

Survey first, then original—surveys are the map, originals are the terrain. The value of a survey: ① gives you a complete coordinate system for a direction in a few pages—what came first, what someone improved, where the field is stuck; ② helps you build a terminology system so you're no longer tripped by unfamiliar concepts when reading originals; ③ the survey's reference list itself is a high-quality paper list. But the downside of surveys is lack of depth and possible lag—post-2020, many directions change so fast that surveys are already outdated before they're finished. So the right posture is: for a new field, first read a recent two-year survey to build a map, then immediately pivot to the two or three originals you care about. Search for surveys on arXiv using "survey" + topic keywords, or follow surveys mentioned in Frontier Progress.

Q12. How many papers until you've "entered the field"? ​

There's no fixed number, but there are verifiable milestones. A practical criterion: when you can draw, without any materials, a structural map of "problem → key methods → mainstream routes → current bottlenecks" for a subfield, and point out three+ representative papers and their relationships, that field is considered "entered." Based on this site's recommendations, the entry path is roughly: nine classics (Perceptron to GPT-3) to build a timeline → pick one deep direction of interest, deep-read 5–10 papers → follow the latest work of that direction from the past two years. Referencing Classic Paper Deep Dives and Paper Map, about 15–30 high-quality papers in one direction can support a level of "being able to follow academic discussion." The entry standard is "can converse," not "has read everything."

3. Summary: Paper Reading Is Long-Termism ​

Viewed on the timescale of time, the conclusion is clear:

Short-term (weeks):  read a few, can't understand, doubt yourself  →  This is every person's rite of passage
Mid-term (months):   deep-read 20 papers + take notes + retell  →  Can clearly keep up with discussions
Long-term (1–2 years): 50–100 papers in one direction  →  Can judge directions independently, propose questions

The compound interest of paper reading shows up in three places: speed—the first Transformer took a week, the tenth attention variant takes an hour; patterned reading gets faster and faster; judgment—the more you read, the faster you spot weak baselines and fancy rhetoric; the red-flag signal table internalizes into intuition; expressiveness—the ability to explain papers clearly is your scarce expressiveness in interviews, reviews, and collaboration.

Finally, back to this FAQ's starting point:

Three disciplines, the three most-worth-taking-away sentences from this article

  1. Take mechanisms, not numbers—every score is a conspiracy of model + data + compute + protocol.
  2. Not understanding is the default state—swap materials, fill prerequisites, mark and re-read; don't go to war with one paper.
  3. Notes are external memory, retelling is the only test—reading without notes is like not reading; understanding you can't explain away is like not understanding.

Can't finish papers, can't understand them, forget them after reading—none of these are problems. There's only one problem: whether you absorb a little reusable thing from every reading session. If yes, it's long-termism; if no, reading ten thousand papers is just writing on sand.

Further Reading ​

  • Start Here — Main entry for the papers section and three entry routes
  • Reading Paths — Specific paper lists and reading order organized by goal
  • Paper Map — Timeline coordinate system of 60 years of key ML papers
  • Classic Paper Deep Dives — Step-by-step deep reads of nine field-changing papers, doubling as note templates
  • Frontier Progress — Tracking 2020s frontiers: large models, diffusion models, Agents, etc.
  • Math Primer — Quick-reference table for when paper math gets stuck
  • awesome resource list — Entry point for finding tutorials, datasets, tools, and official code
  • Learning Paths Overview — Plan your long-term learning from the site's perspective

References ​

All below are real, publicly accessible resources for in-depth self-study: