Skip to content

Reading Paths

Quick overview Paper reading routes organized by goal: Path A, a 2-hour crash course reading 3-5 classics; Path B, engineering deployment mastering BatchNorm, Adam, and fine-tuning papers; Path C, research advancement working through derivations and surveys. Each paper includes reading rationale, depth level, and linked pages, plus reading pace recommendations.

Reading Paths ​

One-sentence definition: Reading paths turn "casually browsing papers" into "reading papers by goal" — decide whether your goal is introduction, deployment, or research first, then pick the corresponding paper list and reading depth, avoiding getting lost in the flood of arXiv.

I. Set Your Goal, Then Choose a Path ​

Papers are countless — what to read and how deeply depends entirely on your goal:

PathGoalWhat to ReadDepth Requirement
A: 2-Hour Crash CourseBuild intuition, follow discussions3-5 classicsUnderstand the ideas, remember the mechanisms
B: Engineering DeploymentTrain reliable modelsTraining recipes, architecture variants, fine-tuningCan explain key designs, ideally reproduce
C: Doing ResearchUnderstand derivations, propose new ideasOriginal derivations, surveys, first-author reproductionCan derive, critique, and extend

The three paths aren't mutually exclusive: A is a prerequisite for B and C, and problems encountered in B will push you toward C. This page is just navigation — the close reading of each paper is in Classic Papers Deep Dive, the global timeline is in Paper Map, the latest developments are in Frontier Advances, and methodology questions (what to do when you can't understand, how to take notes) are in Reading Discipline & FAQ.

II. Path A: 2-Hour Crash Course ​

The goal is "come out of this feeling grounded" — knowing what the most important faces of deep learning look like. We recommend reading in chronological order: first "see the picture, then read the text" — read A Brief History of Deep Learning to build a sense of era, then pick from the 5 papers below, with the first 4 being mandatory.

A1. LeNet (LeCun 1998) — Warm-up, 30 Minutes ​

Why read it: The thinnest entry point. It's the first to fully present the modern CNN skeleton of "convolution → pooling → fully connected → backpropagation training." Short, with intuitive diagrams — you can grasp the overall flow in minutes. Reading depth: Look at the architecture diagram and training flow; no need to dig into formulas. Linked page: Structural details at CNN and Computer Vision.

A2. AlexNet (Krizhevsky 2012) — Close Read, 1-2 Hours ​

Why read it: The landmark event of deep learning's first year. It halved ImageNet error rates using the combo of "ReLU + Dropout + data augmentation + multi-GPU + large ImageNet dataset," and was the first to prove "deep really works." Reading it helps you understand: deep learning is a conspiracy of data, compute, and algorithms. Reading depth: Focus on "why ReLU," "why Dropout," and the experimental comparison tables. Linked page: Activation functions and network structure at Neural Network Fundamentals.

A3. ResNet (He 2016) — Close Read, 1-2 Hours ​

Why read it: A one-sentence contribution that changed a field — "deeper isn't necessarily worse; residual connections make 152-layer networks trainable." It introduced one of the most universal mechanisms in deep learning: identity skip connections. Every large model today (every layer of every Transformer) has ResNet's bones. Reading depth: Fully digest the mathematical form of the residual block $y = F(x) + x$ and the analysis of "why degradation isn't overfitting." Linked page: Network depth and structural design at CNN and Computer Vision.

A4. Attention Is All You Need (Vaswani 2017) — Close Read, 2-3 Hours ​

Why read it: Essential reading among essential reading. The Transformer replaced recurrence with self-attention, ushering in the era of large models. If you don't understand it fully at first, focus on three things: what Q/K/V are, why divide by $\sqrt{d_k}$, and why self-attention is parallelizable. Reading depth: The formulas in the first 3.2 sections are worth going through line by line; skim the rest. Linked page: Full breakdown at Transformer Architecture, mathematical prerequisites at Attention Mechanism.

A5. (Optional) BERT or GPT Papers — Browse, 1 Hour ​

Why read it: Read one representative paper on "pre-training after Transformer" to see how architecture gets activated by training objectives. BERT and GPT are two uses of the same Transformer — understanding one means you understand the other. Reading depth: Browse pre-training tasks and fine-tuning flows; no need to read the appendix. Linked page: Pre-training and fine-tuning panorama at Representation Learning and Pre-training; scaling consequences at Large Language Models (LLM).

Path A Completion Criteria

After closing the book, you should be able to describe in one sentence what each paper "solved and how," and explain to others "why ResNet can train deeper by just adding a shortcut." Once you meet this, proceed to Path B.

III. Path B: Engineering Deployment ​

The goal is "readable and usable" — turn validated training recipes from papers into your default operations. These papers share a characteristic: you can improve your training script the very next day after reading them.

B1. BatchNorm (Ioffe & Szegedy 2015) — Close Read, 1-2 Hours ​

Why read it: It makes "boldly using large learning rates" possible. Understanding why normalizing "per mini-batch along the channel dimension" stabilizes deep training, and why inference switches to global statistics, is a prerequisite for debugging training divergence. Linked page: Comparison with LayerNorm at Initialization and Normalization.

B2. Adam (Kingma & Ba 2015) — Close Read, 1-2 Hours ​

Why read it: The default optimizer in industry. After understanding its three mechanisms (momentum, adaptive learning rate, bias correction), you'll see why "Adam works well by default, but SGD+momentum sometimes converges to flatter, better minima." Linked page: Full landscape of optimization algorithms at Optimization and Gradient Descent.

B3. Transformer Variants: BERT / GPT Series — Optional, 2-3 Hours ​

Why read it: From "one architecture" to "a family of systems." Understanding the Encoder/Decoder distinction, Pre-Norm/Post-Norm, and the evolution of positional encodings helps you make the right choices when fine-tuning pre-trained models (which layers to freeze, what sequence length to use). Linked page: Fine-tuning and parameter-efficient fine-tuning (LoRA, etc.) at Large Language Models (LLM).

B4. Fine-tuning & Alignment: InstructGPT / LoRA / RAG — Optional, 2-3 Hours ​

Why read it: The focus of engineering deployment has shifted from "training from scratch" to "fine-tuning / retrieval-augmented generation / alignment." Understanding the division of labor between RLHF and SFT, why LoRA saves 99% of trainable parameters, and how retrieval-augmented generation fills factual gaps are essential skills for doing LLM applications today. Linked page: Complete methodology for evaluating fine-tuning results at Deep Learning Evaluation and Experiments; training recipe overview at Training Recipes and Hyperparameter Tuning.

B5. (Ongoing) Loss Functions and Regularization Papers — Fragmented Close Reading ​

Why read it: Original papers on cross-entropy, label smoothing, Dropout, and weight decay are all short — great for fragmented reading time. Understanding "why Dropout is equivalent to ensembling" and "why label smoothing suppresses overconfidence" frees you from treating regularization terms as magic spells. Linked page: See Loss Functions and Output Layers and Overfitting and Regularization.

Path B Reminder

After reading these papers, please verify by hand with Training Recipes and Hyperparameter Tuning: move a BatchNorm layer, swap an optimizer, observe the loss curve. If you haven't changed the code yourself, you haven't really read it.

IV. Path C: Research Advancement ​

The goal is "read deeply, explain clearly" — understand derivations, be able to critique, reproduce, and pose your own questions. For those preparing for graduate school, research job interviews, or just wanting to go deep.

C1. Mathematical Derivations: Backpropagation, Adam, Scaling Laws ​

  • Rumelhart et al. 1986 (Backpropagation): Manually derive the chain rule for a three-layer network, understanding the mechanism of "gradient flowing from output to every layer." Fill in math prerequisites with Math Primer.
  • Adam's original derivation: Go through momentum and second-moment bias correction line by line, experiencing "every choice in a derivation has a motivation."
  • Scaling Laws (Kaplan 2020 / Chinchilla 2022): Understand how power-law curves emerge from fitting data, and the meaning of compute-optimal. It turns "how big should the model be" from armchair speculation into a budgeting problem.
  • Two foundational papers on attention: Bahdanau 2015 (attention's origins) and Vaswani 2017 (the Transformer). We recommend manually writing out the matrix dimensions of Q/K/V yourself.

C2. Surveys: Standing on the Shoulders of Giants ​

  • LeCun, Bengio, Hinton 2015 "Deep Learning" (Nature survey): A panoramic map from three founders, extremely high information density, worth rereading;
  • After "Attention Is All You Need": Surveys like A Survey of Transformers (Lin et al. 2021) are great for quickly understanding variant genealogies;
  • Generative models survey: The progressive relationship from VAE to GAN to diffusion models — see Generative Models and VAE and GAN.

C3. First-Author Reproduction: The Researcher's Litmus Test ​

Pick a medium-sized paper (e.g., a simplified ResNet or BERT) and implement it yourself under the framework of building a deep learning project from scratch. Reproduction exposes every place where you thought you understood but didn't — its value is irreplaceable by any notes.

Path C Advice

A researcher's most important ability isn't "reading a lot" but "asking sharp questions." After reading every paper, write down a "what if..." hypothesis, and try to find inspiration in the ablation discussions of Classic Papers Deep Dive.

V. Reading Pace Recommendations ​

TimeRecommended ActionDeliverable
Day 12 papers from Path A1-page notes per paper (question / contribution / ablation)
Week 1Finish remaining Path A + first 2 of Path B30-min close read per paper + 5-min retelling
Weeks 2-3Last 3 of Path B + one hands-on reproductionModify code and compare metrics
Week 4+Path C: derivations + surveys + reproduction1 deep read per week + 1 hypothesis

Rhythm philosophy: 45 minutes of focused reading beats sitting all afternoon; each paper is worth reading three times — first for structure, second for details, third with a critical eye. Put down what you can't understand overnight and reread it the next day — it often clicks (methods detailed in Reading Discipline & FAQ).

VI. Trade-offs ​

  • Breadth vs. depth: Path A aims for breadth, C aims for depth. During introduction, pursue breadth; during research, abandon the vanity of "knowing a little about everything" and dig into one point.
  • Original papers vs. second-hand interpretations: blogs and courses can help you "preview," but always return to the original — second-hand interpretations lose critical details and sometimes transmit wrong conclusions.
  • Classics vs. frontier: 80% of time on classics (stable mechanisms, well-validated), 20% on frontier (see Frontier Advances). Reversing this leads to an unstable foundation.
  • Reading vs. doing: "Feeling like you understand" is the most dangerous illusion. After every group of papers, force yourself to do an experiment in the style of a portfolio project and turn understanding into code.

Further Reading ​

References ​