Theme
Reading Paths
One-sentence definition: Reading paths turn "casually browsing papers" into "reading papers by goal" — decide whether your goal is introduction, deployment, or research first, then pick the corresponding paper list and reading depth, avoiding getting lost in the flood of arXiv.
I. Set Your Goal, Then Choose a Path
Papers are countless — what to read and how deeply depends entirely on your goal:
| Path | Goal | What to Read | Depth Requirement |
|---|---|---|---|
| A: 2-Hour Crash Course | Build intuition, follow discussions | 3-5 classics | Understand the ideas, remember the mechanisms |
| B: Engineering Deployment | Train reliable models | Training recipes, architecture variants, fine-tuning | Can explain key designs, ideally reproduce |
| C: Doing Research | Understand derivations, propose new ideas | Original derivations, surveys, first-author reproduction | Can derive, critique, and extend |
The three paths aren't mutually exclusive: A is a prerequisite for B and C, and problems encountered in B will push you toward C. This page is just navigation — the close reading of each paper is in Classic Papers Deep Dive, the global timeline is in Paper Map, the latest developments are in Frontier Advances, and methodology questions (what to do when you can't understand, how to take notes) are in Reading Discipline & FAQ.
II. Path A: 2-Hour Crash Course
The goal is "come out of this feeling grounded" — knowing what the most important faces of deep learning look like. We recommend reading in chronological order: first "see the picture, then read the text" — read A Brief History of Deep Learning to build a sense of era, then pick from the 5 papers below, with the first 4 being mandatory.
A1. LeNet (LeCun 1998) — Warm-up, 30 Minutes
Why read it: The thinnest entry point. It's the first to fully present the modern CNN skeleton of "convolution → pooling → fully connected → backpropagation training." Short, with intuitive diagrams — you can grasp the overall flow in minutes. Reading depth: Look at the architecture diagram and training flow; no need to dig into formulas. Linked page: Structural details at CNN and Computer Vision.
A2. AlexNet (Krizhevsky 2012) — Close Read, 1-2 Hours
Why read it: The landmark event of deep learning's first year. It halved ImageNet error rates using the combo of "ReLU + Dropout + data augmentation + multi-GPU + large ImageNet dataset," and was the first to prove "deep really works." Reading it helps you understand: deep learning is a conspiracy of data, compute, and algorithms. Reading depth: Focus on "why ReLU," "why Dropout," and the experimental comparison tables. Linked page: Activation functions and network structure at Neural Network Fundamentals.
A3. ResNet (He 2016) — Close Read, 1-2 Hours
Why read it: A one-sentence contribution that changed a field — "deeper isn't necessarily worse; residual connections make 152-layer networks trainable." It introduced one of the most universal mechanisms in deep learning: identity skip connections. Every large model today (every layer of every Transformer) has ResNet's bones. Reading depth: Fully digest the mathematical form of the residual block $y = F(x) + x$ and the analysis of "why degradation isn't overfitting." Linked page: Network depth and structural design at CNN and Computer Vision.
A4. Attention Is All You Need (Vaswani 2017) — Close Read, 2-3 Hours
Why read it: Essential reading among essential reading. The Transformer replaced recurrence with self-attention, ushering in the era of large models. If you don't understand it fully at first, focus on three things: what Q/K/V are, why divide by $\sqrt{d_k}$, and why self-attention is parallelizable. Reading depth: The formulas in the first 3.2 sections are worth going through line by line; skim the rest. Linked page: Full breakdown at Transformer Architecture, mathematical prerequisites at Attention Mechanism.
A5. (Optional) BERT or GPT Papers — Browse, 1 Hour
Why read it: Read one representative paper on "pre-training after Transformer" to see how architecture gets activated by training objectives. BERT and GPT are two uses of the same Transformer — understanding one means you understand the other. Reading depth: Browse pre-training tasks and fine-tuning flows; no need to read the appendix. Linked page: Pre-training and fine-tuning panorama at Representation Learning and Pre-training; scaling consequences at Large Language Models (LLM).
Path A Completion Criteria
After closing the book, you should be able to describe in one sentence what each paper "solved and how," and explain to others "why ResNet can train deeper by just adding a shortcut." Once you meet this, proceed to Path B.
III. Path B: Engineering Deployment
The goal is "readable and usable" — turn validated training recipes from papers into your default operations. These papers share a characteristic: you can improve your training script the very next day after reading them.
B1. BatchNorm (Ioffe & Szegedy 2015) — Close Read, 1-2 Hours
Why read it: It makes "boldly using large learning rates" possible. Understanding why normalizing "per mini-batch along the channel dimension" stabilizes deep training, and why inference switches to global statistics, is a prerequisite for debugging training divergence. Linked page: Comparison with LayerNorm at Initialization and Normalization.
B2. Adam (Kingma & Ba 2015) — Close Read, 1-2 Hours
Why read it: The default optimizer in industry. After understanding its three mechanisms (momentum, adaptive learning rate, bias correction), you'll see why "Adam works well by default, but SGD+momentum sometimes converges to flatter, better minima." Linked page: Full landscape of optimization algorithms at Optimization and Gradient Descent.
B3. Transformer Variants: BERT / GPT Series — Optional, 2-3 Hours
Why read it: From "one architecture" to "a family of systems." Understanding the Encoder/Decoder distinction, Pre-Norm/Post-Norm, and the evolution of positional encodings helps you make the right choices when fine-tuning pre-trained models (which layers to freeze, what sequence length to use). Linked page: Fine-tuning and parameter-efficient fine-tuning (LoRA, etc.) at Large Language Models (LLM).
B4. Fine-tuning & Alignment: InstructGPT / LoRA / RAG — Optional, 2-3 Hours
Why read it: The focus of engineering deployment has shifted from "training from scratch" to "fine-tuning / retrieval-augmented generation / alignment." Understanding the division of labor between RLHF and SFT, why LoRA saves 99% of trainable parameters, and how retrieval-augmented generation fills factual gaps are essential skills for doing LLM applications today. Linked page: Complete methodology for evaluating fine-tuning results at Deep Learning Evaluation and Experiments; training recipe overview at Training Recipes and Hyperparameter Tuning.
B5. (Ongoing) Loss Functions and Regularization Papers — Fragmented Close Reading
Why read it: Original papers on cross-entropy, label smoothing, Dropout, and weight decay are all short — great for fragmented reading time. Understanding "why Dropout is equivalent to ensembling" and "why label smoothing suppresses overconfidence" frees you from treating regularization terms as magic spells. Linked page: See Loss Functions and Output Layers and Overfitting and Regularization.
Path B Reminder
After reading these papers, please verify by hand with Training Recipes and Hyperparameter Tuning: move a BatchNorm layer, swap an optimizer, observe the loss curve. If you haven't changed the code yourself, you haven't really read it.
IV. Path C: Research Advancement
The goal is "read deeply, explain clearly" — understand derivations, be able to critique, reproduce, and pose your own questions. For those preparing for graduate school, research job interviews, or just wanting to go deep.
C1. Mathematical Derivations: Backpropagation, Adam, Scaling Laws
- Rumelhart et al. 1986 (Backpropagation): Manually derive the chain rule for a three-layer network, understanding the mechanism of "gradient flowing from output to every layer." Fill in math prerequisites with Math Primer.
- Adam's original derivation: Go through momentum and second-moment bias correction line by line, experiencing "every choice in a derivation has a motivation."
- Scaling Laws (Kaplan 2020 / Chinchilla 2022): Understand how power-law curves emerge from fitting data, and the meaning of compute-optimal. It turns "how big should the model be" from armchair speculation into a budgeting problem.
- Two foundational papers on attention: Bahdanau 2015 (attention's origins) and Vaswani 2017 (the Transformer). We recommend manually writing out the matrix dimensions of Q/K/V yourself.
C2. Surveys: Standing on the Shoulders of Giants
- LeCun, Bengio, Hinton 2015 "Deep Learning" (Nature survey): A panoramic map from three founders, extremely high information density, worth rereading;
- After "Attention Is All You Need": Surveys like A Survey of Transformers (Lin et al. 2021) are great for quickly understanding variant genealogies;
- Generative models survey: The progressive relationship from VAE to GAN to diffusion models — see Generative Models and VAE and GAN.
C3. First-Author Reproduction: The Researcher's Litmus Test
Pick a medium-sized paper (e.g., a simplified ResNet or BERT) and implement it yourself under the framework of building a deep learning project from scratch. Reproduction exposes every place where you thought you understood but didn't — its value is irreplaceable by any notes.
Path C Advice
A researcher's most important ability isn't "reading a lot" but "asking sharp questions." After reading every paper, write down a "what if..." hypothesis, and try to find inspiration in the ablation discussions of Classic Papers Deep Dive.
V. Reading Pace Recommendations
| Time | Recommended Action | Deliverable |
|---|---|---|
| Day 1 | 2 papers from Path A | 1-page notes per paper (question / contribution / ablation) |
| Week 1 | Finish remaining Path A + first 2 of Path B | 30-min close read per paper + 5-min retelling |
| Weeks 2-3 | Last 3 of Path B + one hands-on reproduction | Modify code and compare metrics |
| Week 4+ | Path C: derivations + surveys + reproduction | 1 deep read per week + 1 hypothesis |
Rhythm philosophy: 45 minutes of focused reading beats sitting all afternoon; each paper is worth reading three times — first for structure, second for details, third with a critical eye. Put down what you can't understand overnight and reread it the next day — it often clicks (methods detailed in Reading Discipline & FAQ).
VI. Trade-offs
- Breadth vs. depth: Path A aims for breadth, C aims for depth. During introduction, pursue breadth; during research, abandon the vanity of "knowing a little about everything" and dig into one point.
- Original papers vs. second-hand interpretations: blogs and courses can help you "preview," but always return to the original — second-hand interpretations lose critical details and sometimes transmit wrong conclusions.
- Classics vs. frontier: 80% of time on classics (stable mechanisms, well-validated), 20% on frontier (see Frontier Advances). Reversing this leads to an unstable foundation.
- Reading vs. doing: "Feeling like you understand" is the most dangerous illusion. After every group of papers, force yourself to do an experiment in the style of a portfolio project and turn understanding into code.
Further Reading
- Learning Paths: Three Routes — the site's overall learning paths and where paper reading fits
- What Is Deep Learning — the minimum conceptual foundation before reading papers
- A Brief History of Deep Learning — the timeline version of the paper map
- Glossary — check here first for unfamiliar terms while reading papers
- Awesome List — paper repos and reading tools
References
- LeCun, Bottou, Bengio, Haffner. Gradient-Based Learning Applied to Document Recognition (Proceedings of the IEEE 1998)
- Krizhevsky, Sutskever, Hinton. ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012)
- He, Zhang, Ren, Sun. Deep Residual Learning for Image Recognition (CVPR 2016)
- Vaswani et al. Attention Is All You Need (NeurIPS 2017)
- Ioffe, Szegedy. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift (ICML 2015)
- Kingma, Ba. Adam: A Method for Stochastic Optimization (ICLR 2015)
- Kaplan et al. Scaling Laws for Neural Language Models (2020)
- Hoffmann et al. Training Compute-Optimal Large Language Models (Chinchilla) (NeurIPS 2022)