Theme
Papers: Start Here
One-sentence definition: Reading papers is the necessary path for deep learning learners to go from "can tune APIs" to "understand principles" — frameworks get replaced, APIs get renamed, but the key ideas that shape the field are written in papers.
I. Why Deep Learning Learners Should Read Papers
1. Understand Principles, Don't Memorize APIs
Most tutorials teach you "how to use it"; papers tell you "why it works." One line of nn.Conv2d in PyTorch is backed by the combo of "convolution + ReLU + Dropout + multi-GPU" validated by AlexNet; optim.Adam synthesizes the ideas of momentum and adaptive learning rates from optimization and gradient descent. People who only know how to call APIs can only search blogs and hope for the best when facing new problems; people who understand papers can reason from first principles about what to use and what not to.
The problem with blogs and WeChat public accounts is: they are second-hand information. Authors restate things through their own understanding, and details get lost, conclusions get simplified, or worse, get transmitted incorrectly. Papers are first-hand information — motivations, assumptions, derivations, failures, and boundary conditions are all in there. The quality of your understanding of a mechanism is capped by the quality of the sources you rely on; if you only read second-hand sources, your understanding will always carry someone else's filter.
Analogy: knowing how to use appliances doesn't mean you understand circuits. You can't fix a broken appliance yourself, but with a circuit diagram, every fault is traceable. The same applies to deep learning — those key papers are the "circuit diagrams" in your hands.
2. Stay on the Frontier
Deep learning advances at a monthly pace: the Transformer appeared in 2017, BERT/GPT scaled up in 2019, diffusion models went mainstream in 2022, and reasoning models redefined "thinking" in 2024. News and blogs only give you conclusions; papers give you complete derivations and experimental evidence. Only after reading papers can you truly understand discussions about MoE, RLHF, diffusion-based video generation, rather than regurgitating second-hand opinions.
More importantly, frontier judgment can only be trained through reading papers. When a new model goes viral, those who have read papers ask "who is the baseline, what ablations were done, is it open-sourced?" while those who haven't just go along with the hype. This gap gets amplified in interviews and at work.
3. Interview Advantage
Algorithm interviews at major tech companies almost always ask about papers: either "talk about the model you're most familiar with" or "how would you design an XX system from scratch?" Being able to explain a paper's "background / method / experiments / limitations" within 5 minutes is a far rarer skill than grinding LeetCode (see Reading Discipline and FAQ for how). Skim through our interview question bank and you'll find the answers to many questions are conclusions from papers.
4. Why Textbooks and Blogs Can't Replace Papers
Textbooks give you "consensus after sedimentation" — they are stable and readable, but they smooth over "controversy" and "why." Papers preserve the scene of exploration: the author's initial ideas, failed attempts, and struggles with baselines are all in there. Reading papers isn't about reading conclusions; it's about reading "how conclusions came to be." For those who truly want to understand deep learning, this "origin story" is exactly what textbooks can't provide.
Key Insight
Reading papers isn't a "mark it as done" checklist task; it's training a way of thinking: breaking any problem down into five dimensions — data, architecture, objective function, optimization, and evaluation. This framework is fully isomorphic with the overall architecture anatomy of deep learning.
II. Three Reading Paths
Different goals call for different paper lists and reading orders. See the "Reading Paths" page for details:
| Path | Who It's For | Goal | Estimated Time |
|---|---|---|---|
| A: 2-Hour Crash Course | Time-constrained, want to build intuition | Close-read 3-5 classics, understand core ideas | 2+ hours |
| B: Engineering Deployment | Engineers who need results | Master training recipes and deployment-related papers | 1-2 weeks |
| C: Doing Research | Graduate students / deeper study / interview prep | Work through derivations, read surveys, implement from scratch | Ongoing, months |
How the three paths connect: almost everyone starts from A — build intuition for what deep learning looks like in the shortest time; if you're already writing training code, you can jump straight to B, but the 2-3 mechanism papers in A are still worth catching up on; if you're aiming for research or algorithm roles, treat C as a long-term habit, not a one-time sprint. The paths aren't mutually exclusive: A is the foundation for B and C, and the bottlenecks you hit in B will naturally push you toward C.
III. The Five Pages of This Section
This section has five pages, each answering one question:
| Page | One-Sentence Summary | Link |
|---|---|---|
| Reading Paths | Choose a paper list and reading order based on your goal | Reading Paths |
| Paper Map | See the full picture of deep learning through themes and a timeline | Paper Map |
| Classic Papers Deep Dive | Close-read a dozen+ classics that changed the field | Classic Papers Deep Dive |
| Frontier Advances | Important breakthroughs and trend analysis from recent years | Frontier Advances |
| Reading Discipline & FAQ | Methodology: paper selection, note-taking, reproduction, interview presentation | Reading Discipline & FAQ |
Suggested order: start with the Paper Map to build a global view, then advance paper by paper following Reading Paths, and finally use Classic Papers Deep Dive for deep work. Check "Reading Discipline & FAQ" for methodology questions whenever needed. In short: map positioning → path planning → close reading execution → methodology support — the four pages each have their role. Save Frontier Advances until you have a solid foundation.
Common Use Cases
- Interview sprint: open the quick-reference table of 13 papers in "Classic Papers Deep Dive," pick the models you mentioned on your resume, close-read them one by one, then practice explaining them using the 5-step method from "Reading Discipline & FAQ";
- Training your own model: go straight to path B in "Reading Paths," treat BatchNorm, Adam, and regularization papers as operation manuals, and immediately tweak your training scripts after reading;
- Overwhelmed by frontier hype: check "what it is, why it matters, representative work" in "Frontier Advances," then follow references to read the originals — skip the second-hand emotion;
- Want to judge whether a paper is good: run through the four criteria in "Reading Discipline & FAQ" (baseline, ablation, open-source, data availability) — more reliable than any leaderboard.
IV. The Most Important Advice: Take Away Mechanisms, Not Numbers
The most easily dated things in papers are numbers: error rates drop year by year, parameter counts double, and training time shrinks dramatically with better hardware. What's truly valuable are mechanisms:
- AlexNet's 15.3% top-5 error rate has long been surpassed, but the mechanism of "depth + big data + strong compute working together" still dominates computer vision;
- GPT-3's 175B parameter count will be surpassed, but the mechanism of "in-context learning" defined the entire era of large language models;
- DDPM's metrics will keep getting updated, but the mechanism of "forward noising + reverse denoising" has been fully inherited by video, 3D, and audio generation.
How to tell "mechanism" from "numbers"? Three criteria:
- Does it hold in a different setting? Mechanisms transfer across tasks and datasets; numbers don't.
- Does it hold at a different scale? Mechanisms work for small models too; numbers change dramatically with scale.
- Does it explain "why"? Mechanisms answer "why it works"; numbers only answer "how well."
After close-reading each paper, ask yourself three questions: What problem does it solve? What is its core contribution? What do the ablation experiments prove? If you can answer these three, it's fine if you can't remember the numbers; if you can't, memorizing numbers won't help (for the specific note-taking framework, see the "Three Questions" section in "Reading Discipline & FAQ").
V. Boundaries and Trade-offs
- Papers aren't the only learning path: for beginners, first run through the progressive tutorial: three versions, get it running to build feel, then read papers — this is far more efficient than "toughing through." The relationship between practice and papers is covered in building a deep learning project from scratch.
- Don't fall into "paper anxiety": thousands of arXiv papers are published daily; you don't need to read them all. Lock in your goal and invest 80% of your energy in classic papers and your sub-field.
- Be cautious with time-sensitive data: model names and leaderboard numbers in the Frontier Advances page expire quickly. Pay attention to the cutoff date notice at the top of that page and cross-check with the continuously updated awesome list.
- Learn terms first: when stuck on a paper, check the glossary to understand the terminology first, then come back to the formulas — it often clicks.
- Don't skip hands-on: reading 100 papers is worth less than reproducing one from scratch. Turn every paper's takeaway into code or experiments, and the knowledge will truly stick with you — this is more important than any reading technique.
One-sentence summary: Papers are the most honest and complete knowledge carrier in this field. Our goal isn't to make you read all papers, but to give you the ability to "understand a paper" — once established, this ability compounds over your entire career.
Further Reading
- Learning Paths: Three Routes — three learning paths at the site level, complementary to this section
- A Brief History of Deep Learning — a domain timeline you can grasp before reading papers
- Interview Question Bank — translating paper knowledge into interview answers
- Portfolio Projects — the best outlet for hands-on validation after reading papers
- Awesome List — continuously updated courses, paper repos, and tools
References
- LeCun, Bengio, Hinton. Deep Learning (Nature 2015) — a domain survey co-written by the three founders of deep learning, essential reading for any learner
- Vaswani et al. Attention Is All You Need (NeurIPS 2017)
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020)
- Ho, Jain, Abbeel. Denoising Diffusion Probabilistic Models (NeurIPS 2020)
- Rosenblatt. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (Psychological Review 1958)