Skip to content

Papers: Start Here

Quick overview Why deep learning learners should read papers: understand principles instead of memorizing APIs, stay on the frontier, and stand out in interviews. This page presents three reading paths (2-hour crash course / engineering deployment / research) and the five-page map of this section, along with the most important advice — take away mechanisms, not numbers.

Papers: Start Here ​

One-sentence definition: Reading papers is the necessary path for deep learning learners to go from "can tune APIs" to "understand principles" — frameworks get replaced, APIs get renamed, but the key ideas that shape the field are written in papers.

I. Why Deep Learning Learners Should Read Papers ​

1. Understand Principles, Don't Memorize APIs ​

Most tutorials teach you "how to use it"; papers tell you "why it works." One line of nn.Conv2d in PyTorch is backed by the combo of "convolution + ReLU + Dropout + multi-GPU" validated by AlexNet; optim.Adam synthesizes the ideas of momentum and adaptive learning rates from optimization and gradient descent. People who only know how to call APIs can only search blogs and hope for the best when facing new problems; people who understand papers can reason from first principles about what to use and what not to.

The problem with blogs and WeChat public accounts is: they are second-hand information. Authors restate things through their own understanding, and details get lost, conclusions get simplified, or worse, get transmitted incorrectly. Papers are first-hand information — motivations, assumptions, derivations, failures, and boundary conditions are all in there. The quality of your understanding of a mechanism is capped by the quality of the sources you rely on; if you only read second-hand sources, your understanding will always carry someone else's filter.

Analogy: knowing how to use appliances doesn't mean you understand circuits. You can't fix a broken appliance yourself, but with a circuit diagram, every fault is traceable. The same applies to deep learning — those key papers are the "circuit diagrams" in your hands.

2. Stay on the Frontier ​

Deep learning advances at a monthly pace: the Transformer appeared in 2017, BERT/GPT scaled up in 2019, diffusion models went mainstream in 2022, and reasoning models redefined "thinking" in 2024. News and blogs only give you conclusions; papers give you complete derivations and experimental evidence. Only after reading papers can you truly understand discussions about MoE, RLHF, diffusion-based video generation, rather than regurgitating second-hand opinions.

More importantly, frontier judgment can only be trained through reading papers. When a new model goes viral, those who have read papers ask "who is the baseline, what ablations were done, is it open-sourced?" while those who haven't just go along with the hype. This gap gets amplified in interviews and at work.

3. Interview Advantage ​

Algorithm interviews at major tech companies almost always ask about papers: either "talk about the model you're most familiar with" or "how would you design an XX system from scratch?" Being able to explain a paper's "background / method / experiments / limitations" within 5 minutes is a far rarer skill than grinding LeetCode (see Reading Discipline and FAQ for how). Skim through our interview question bank and you'll find the answers to many questions are conclusions from papers.

4. Why Textbooks and Blogs Can't Replace Papers ​

Textbooks give you "consensus after sedimentation" — they are stable and readable, but they smooth over "controversy" and "why." Papers preserve the scene of exploration: the author's initial ideas, failed attempts, and struggles with baselines are all in there. Reading papers isn't about reading conclusions; it's about reading "how conclusions came to be." For those who truly want to understand deep learning, this "origin story" is exactly what textbooks can't provide.

Key Insight

Reading papers isn't a "mark it as done" checklist task; it's training a way of thinking: breaking any problem down into five dimensions — data, architecture, objective function, optimization, and evaluation. This framework is fully isomorphic with the overall architecture anatomy of deep learning.

II. Three Reading Paths ​

Different goals call for different paper lists and reading orders. See the "Reading Paths" page for details:

PathWho It's ForGoalEstimated Time
A: 2-Hour Crash CourseTime-constrained, want to build intuitionClose-read 3-5 classics, understand core ideas2+ hours
B: Engineering DeploymentEngineers who need resultsMaster training recipes and deployment-related papers1-2 weeks
C: Doing ResearchGraduate students / deeper study / interview prepWork through derivations, read surveys, implement from scratchOngoing, months

How the three paths connect: almost everyone starts from A — build intuition for what deep learning looks like in the shortest time; if you're already writing training code, you can jump straight to B, but the 2-3 mechanism papers in A are still worth catching up on; if you're aiming for research or algorithm roles, treat C as a long-term habit, not a one-time sprint. The paths aren't mutually exclusive: A is the foundation for B and C, and the bottlenecks you hit in B will naturally push you toward C.

III. The Five Pages of This Section ​

This section has five pages, each answering one question:

PageOne-Sentence SummaryLink
Reading PathsChoose a paper list and reading order based on your goalReading Paths
Paper MapSee the full picture of deep learning through themes and a timelinePaper Map
Classic Papers Deep DiveClose-read a dozen+ classics that changed the fieldClassic Papers Deep Dive
Frontier AdvancesImportant breakthroughs and trend analysis from recent yearsFrontier Advances
Reading Discipline & FAQMethodology: paper selection, note-taking, reproduction, interview presentationReading Discipline & FAQ

Suggested order: start with the Paper Map to build a global view, then advance paper by paper following Reading Paths, and finally use Classic Papers Deep Dive for deep work. Check "Reading Discipline & FAQ" for methodology questions whenever needed. In short: map positioning → path planning → close reading execution → methodology support — the four pages each have their role. Save Frontier Advances until you have a solid foundation.

Common Use Cases ​

  • Interview sprint: open the quick-reference table of 13 papers in "Classic Papers Deep Dive," pick the models you mentioned on your resume, close-read them one by one, then practice explaining them using the 5-step method from "Reading Discipline & FAQ";
  • Training your own model: go straight to path B in "Reading Paths," treat BatchNorm, Adam, and regularization papers as operation manuals, and immediately tweak your training scripts after reading;
  • Overwhelmed by frontier hype: check "what it is, why it matters, representative work" in "Frontier Advances," then follow references to read the originals — skip the second-hand emotion;
  • Want to judge whether a paper is good: run through the four criteria in "Reading Discipline & FAQ" (baseline, ablation, open-source, data availability) — more reliable than any leaderboard.

IV. The Most Important Advice: Take Away Mechanisms, Not Numbers ​

The most easily dated things in papers are numbers: error rates drop year by year, parameter counts double, and training time shrinks dramatically with better hardware. What's truly valuable are mechanisms:

  • AlexNet's 15.3% top-5 error rate has long been surpassed, but the mechanism of "depth + big data + strong compute working together" still dominates computer vision;
  • GPT-3's 175B parameter count will be surpassed, but the mechanism of "in-context learning" defined the entire era of large language models;
  • DDPM's metrics will keep getting updated, but the mechanism of "forward noising + reverse denoising" has been fully inherited by video, 3D, and audio generation.

How to tell "mechanism" from "numbers"? Three criteria:

  1. Does it hold in a different setting? Mechanisms transfer across tasks and datasets; numbers don't.
  2. Does it hold at a different scale? Mechanisms work for small models too; numbers change dramatically with scale.
  3. Does it explain "why"? Mechanisms answer "why it works"; numbers only answer "how well."

After close-reading each paper, ask yourself three questions: What problem does it solve? What is its core contribution? What do the ablation experiments prove? If you can answer these three, it's fine if you can't remember the numbers; if you can't, memorizing numbers won't help (for the specific note-taking framework, see the "Three Questions" section in "Reading Discipline & FAQ").

V. Boundaries and Trade-offs ​

  • Papers aren't the only learning path: for beginners, first run through the progressive tutorial: three versions, get it running to build feel, then read papers — this is far more efficient than "toughing through." The relationship between practice and papers is covered in building a deep learning project from scratch.
  • Don't fall into "paper anxiety": thousands of arXiv papers are published daily; you don't need to read them all. Lock in your goal and invest 80% of your energy in classic papers and your sub-field.
  • Be cautious with time-sensitive data: model names and leaderboard numbers in the Frontier Advances page expire quickly. Pay attention to the cutoff date notice at the top of that page and cross-check with the continuously updated awesome list.
  • Learn terms first: when stuck on a paper, check the glossary to understand the terminology first, then come back to the formulas — it often clicks.
  • Don't skip hands-on: reading 100 papers is worth less than reproducing one from scratch. Turn every paper's takeaway into code or experiments, and the knowledge will truly stick with you — this is more important than any reading technique.

One-sentence summary: Papers are the most honest and complete knowledge carrier in this field. Our goal isn't to make you read all papers, but to give you the ability to "understand a paper" — once established, this ability compounds over your entire career.

Further Reading ​

References ​