Theme
Reading Paths
In one sentence: Reading papers isn't about "reading a lot," it's about having the "right path"—first set your goal, then decide which papers to read and how deeply. This article gives you three curated roadmaps, plus a universal method for fully understanding any paper.
1. Why Read Along a Path: Navigating the Sea of Papers
Consider these numbers: on arXiv alone, the cs.LG (Machine Learning) category adds over 20,000 new papers every year; add cs.CV (Computer Vision), cs.CL (Computational Linguistics/NLP), cs.AI, and stat.ML, and hundreds of new papers arrive daily. Including historical accumulation, the total number of ML papers already numbers in the millions.
In the face of this scale, two brutal facts emerge:
- No one can read them all. Even top researchers in a field can only deeply read three to five papers and skim dozens per week; the rest are filtered by title, abstract, and secondary curation.
- Not filtering is a disaster. Paper quality is extremely uneven: there are paradigm-shifting foundational works, engineering tuning reports, and papers published solely to pad publication counts. Reading without selection will leave you with a pile of disconnected fragments three months later, unable to articulate the skeleton of the field.
This is what "path" means. Path = Goal + Sequence + Depth Standard:
No path: Daily arXiv scroll → read diffusion today, Agent tomorrow
→ Three months later: seen everything, can't explain anything
With path: Set goal (job hunting / production / research)
→ Choose route (quick start / systematic / frontier)
→ Read in curated order (each paper builds on the previous)
→ Three months later: a structured mental mapFollowing a path, every paper you read is not an isolated document but a node on a map—the question raised by the previous paper is exactly answered by the next. This "interlocking knowledge nodes" effect gives your understanding compound interest.
How paths relate to the map, deep dive, and frontier
The papers section of this site is a cohesive set: the Paper Map answers "what are the key papers across 60 years, and who came first?" (macro coordinate system); Classic Paper Deep Dives answers "how to fully absorb a paper" (microscope); Frontier Progress answers "where is the field heading recently" (radar). The Reading Paths here ties these three into an action plan: read in order, choose based on goal, persist at a steady pace.
2. Overview of Three Paths
Based on learners' core goals, this site organizes paper reading into three paths. They are not "ranks" but "directions":
| Path | Core Goal | Time Budget | Who It's For | End Goal |
|---|---|---|---|---|
| ① Quick Start | Build confidence and a big-picture sense that "I can read a paper" | ~4 hours | Beginners who've just learned foundational concepts and want to touch papers for the first time | Deeply read 4 classics, able to explain "why large models work" |
| ② Systematic Study | Fully internalize the skeleton of deep learning | ~2 weeks (1 group per week) | Job seekers, engineers, grad students planning deep theoretical study | Mastery of mechanisms + ablation logic + interview talking points across four themes |
| ③ Frontier Tracking | Keep up with the latest trends of the 2020s | Ongoing (1–2 hours per week) | Advanced learners already in industry or academia | A steady information pipeline: know how to find, filter, and judge |
Quick Start ── 4 hours ──► Can read: build confidence & big-picture sense (for first-time paper readers)
Systematic Study ── 2 weeks ──► Fully absorb: complete skeleton + mechanisms + talking points (for job hunting / research)
Frontier Tracking ── Ongoing ──► Keep up: find, filter, and judge trends (for continuous advancement)These three paths are not "pick one"—they are three stages
The Quick Start path is the starting point for everyone, including grad students who will later go into research. Spend 4 hours first to build the confidence of "I can understand a paper," then decide whether to spend 2 weeks on systematic study, and finally make frontier tracking a daily habit. Skipping the first stage and diving straight into the latest arXiv papers is what causes most people to get "discouraged from reading papers."
3. Path One: Quick Start (~4 hours)
1. Route Design: Why These Four Papers
The Quick Start path has only 4 papers. The order is deliberately not chronological but based on "cognitive dependency": first build a concrete picture of "what large models can actually do" with ReAct, then go back and fill in the principles.
ReAct (2022) ── Phenomenon: large models can "think" and "act," far surpassing expectations
│
Attention Is All You Need (2017) ── Principle: the capability comes from self-attention
│
BERT (2018) ── Paradigm: pretraining + fine-tuning, unifying tasks
│
GPT-3 (2020) ── Pattern: scale is power, emerging few-shot capability2. One Sentence Per Paper
| Order | Paper | Why Read It in One Sentence |
|---|---|---|
| 1 | ReAct (Yao et al., 2022) | It demonstrates how "large models + reasoning + tool use" becomes an Agent capable of autonomous action—a jaw-dropping opening that gives you an "explainable to anyone" wow moment, building interest and big-picture sense |
| 2 | Attention Is All You Need (Vaswani et al., 2017) | It's the birth certificate of the Transformer, answering what the "engine" behind ReAct's models is—self-attention. Without it, there is no BERT, no GPT, no Agent |
| 3 | BERT (Devlin et al., 2018) | It established the "pretraining + fine-tuning" paradigm that dominated the 2020s—340M parameters sweeping 11 NLP benchmarks, the best specimen for "why pretraining works" |
| 4 | GPT-3 (Brown et al., 2020) | It pushed the "scaling law" to its extreme: at 175 billion parameters, the model completes tasks it was never specifically trained on, relying only on prompts (few-shot)—the origin of "why we scale parameters" |
3. Execution Pacing
- Hour 1: Read ReAct. No need to read every word; focus on the "Thought → Action → Observation" loop it designs. Pair with a Chinese visual explanation (see the site's Frontier Progress overview on Agents).
- Hour 2: Read the abstract of Attention Is All You Need, Figure 1 (encoder-decoder structure), and the multi-head attention formula (Section 3.2), paired with Jay Alammar's The Illustrated Transformer.
- Hour 3: Read BERT's abstract, its pretraining tasks (MLM + NSP), and the model structure in Figure 1. You only need to grasp one sentence: use massive unlabeled text for cloze-fill and next-sentence prediction, get a universal representation, then fine-tune for downstream tasks.
- Hour 4: Read GPT-3's abstract, the few-shot setup in Section 3, and the experiment result curves. Internalize one intuition: parameters × data × compute are the three knobs of language intelligence.
Success criteria for the Quick Start path
After reading 4 papers, you should be able to without looking at notes explain to someone else three things: (1) large models can act autonomously through "reasoning + tools" (ReAct); (2) they process information in parallel via self-attention (Transformer); (3) pretraining gives models universal capabilities, and scale gives them emergent few-shot abilities (BERT → GPT-3). If you can explain these three, your first paper-reading experience is a success.
A deeper version of these four
The Quick Start path is a "tasting" session—reading only the abstract, introduction, figures, and conclusions of each paper, totaling about 4 hours. When you're ready to deep-dive them into the Systematic Study Path, use the three-pass method in Section 6 to fully absorb each one; the site's Classic Paper Deep Dives already has step-by-step breakdowns of Attention, BERT, and GPT-3 that you can cross-reference directly.
4. Path Two: Systematic Study (~2 weeks)
1. Route Design: Four Theme Groups
The Systematic Study path divides deep learning into four theme groups, one group per week, 2–3 papers per group, totaling about 2 weeks. The grouping logic follows the natural evolution of the field: first perception (CNN), then capability (attention), then creation (generation), then decision-making (reinforcement learning).
Week 1 Classical Foundations AlexNet → ResNet "Why deep learning works"
Week 2 Attention & Pretraining Attention → BERT → GPT-3 "The engine of language intelligence"
Week 3 Generative Models GAN → Diffusion "How machines create"
Week 4 Reinforcement & Decision AlphaGo "How machines make sequential decisions"2. Group Arrangements and Verification Criteria
Week 1: Classical Foundations (AlexNet → ResNet)
| Paper | Year | Core Contribution |
|---|---|---|
| AlexNet | 2012 | 8-layer CNN + ReLU + Dropout + data augmentation + GPU parallelism, reduced ImageNet top-5 error from 26.2% to 15.3%, igniting the deep learning revolution |
| ResNet | 2015 | 152-layer residual network, solving gradient degradation in deep networks via "identity shortcut connections," first exceeding human-level performance on ImageNet (top-5 error 3.57%) |
- Key takeaway: The three design principles of CNNs—local connectivity, weight sharing, hierarchical features; the residual connection mechanism of "giving gradients a highway to travel."
- Verification: Draw a ResNet residual block from memory and explain "why adding an identity shortcut makes training 152 layers possible."
Week 2: Attention & Pretraining (Attention → BERT → GPT-3)
| Paper | Year | Core Contribution |
|---|---|---|
| Attention Is All You Need | 2017 | Replaced RNN with pure self-attention, proposed multi-head attention and positional encoding—parallelizable training, leap in long-range dependency handling |
| BERT | 2018 | Bidirectional Transformer + Masked Language Model (MLM), established the "pretraining + fine-tuning" paradigm |
| GPT-3 | 2020 | 175 billion parameters + massive data, demonstrated in-context learning and few-shot emergence |
- Key takeaway: The geometric intuition of the Q/K/V attention formula
Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V; the division of labor between "pretraining learns the general, fine-tuning learns the specific"; the phenomenon of capability leaps as scale grows (emergence). - Verification: Derive and explain the computation flow of multi-head attention on a whiteboard; clearly articulate the essential difference between BERT's bidirectional and GPT's autoregressive pretraining objectives.
Week 3: Generative Models (GAN → Diffusion)
| Paper | Year | Core Contribution |
|---|---|---|
| GAN | 2014 | Adversarial training of generator and discriminator, enabling generative models to produce realistic images for the first time (a game of "forgery and detection") |
| DDPM (Diffusion Model) | 2020 | Forward noise-adding, reverse noise-removing, simple and stable training objective, became the SOTA foundation for image generation after 2022 |
- Key takeaway: The meaning of the adversarial training objective function
min max; the Markov chain intuition of diffusion models ("forward noise-adding + reverse noise-removing") and why diffusion is more stable to train than GAN. - Verification: Draw the training loop diagrams for GAN and diffusion models respectively, and name one reason for instability/stability for each.
Week 4: Reinforcement Learning & Decision (AlphaGo)
| Paper | Year | Core Contribution |
|---|---|---|
| AlphaGo | 2016 | Combined policy network + value network + Monte Carlo Tree Search (MCTS), defeated a human world champion at Go—the pinnacle demonstration of "deep learning + RL" |
- Key takeaway: RL's "state → action → reward" loop; how the policy network (predicting the next move) and value network (evaluating the current position) complement each other; how MCTS combines "deep learning intuition" with "search precision."
- Verification: Explain "why AlphaGo couldn't use pure search or pure neural networks, but needed both."
3. Execution Discipline for the Systematic Study Path
- Set aside a fixed 1–2 hours daily; don't cram. Leave one day between groups to "close the notebook and recite from memory."
- Every paper must produce notes: a three-sentence summary + one mechanism diagram + one "what if we removed this design?" counterquestion. See Reading Discipline & FAQ for note-taking methods.
- After finishing each group, do a review: what is the internal evolution logic of this theme? What leftover problem from the previous paper does the next paper solve?
Two hurdles you may encounter on the Systematic Study Path
First, math discouragement — you may get stuck when hitting attention formulas or diffusion model objective functions. Countermeasure: get the intuition first (what is this paper trying to do?), then the derivation (why designed this way?), and only then drill into proof details. Stopping at the second layer is enough for most readers. Second, the temptation to reproduce — wanting to immediately run the paper's code. Countermeasure: the systematic study phase prioritizes understanding; save hands-on reproduction for after you've finished. See Resources for paper code aggregations.
5. Path Three: Frontier Tracking (Ongoing)
The Systematic Study path gives you the "stock" skeleton; the Frontier Tracking path handles the "increment"—consistently integrating new papers into your knowledge map. Frontier tracking is not "daily scrolling" (that only creates anxiety); it's a pipeline:
Information sources (arXiv / top conferences / leaderboards)
↓ filtered
Candidate list (10–20 papers per week)
↓ prioritized
Worth reading (3–5 per week) → skim 2 / deep-read 1
↓ solidify
One-sentence cards filed (what problem does this solve? what existing knowledge does it connect to?)1. Primary Information Source: arXiv
arXiv is the primary publishing venue for all ML papers—almost all new work appears as a preprint first. Usage:
- Subscribe by category:
cs.LG(Machine Learning),cs.CV(Computer Vision),cs.CL(Natural Language),stat.ML,cs.AI. After daily updates, just skim titles—only click on the abstract when a title catches your attention. - Follow specific authors: Every direction has 5–10 core authors. Bookmarking their pages and periodically scanning for new submissions is far more efficient than scrolling entire categories daily.
- Use citation tracking: After reading a good paper, check "papers that cite it" (Google Scholar's Cited by), which automatically moves you forward along the timeline.
2. Leaderboards & Reproduction: Papers with Code
Papers with Code ties papers, code, and evaluation benchmarks (SOTA leaderboards) together. Three ways to use it:
- Check "currently best": Look at the SOTA leaderboard for a task (e.g., ImageNet, GLUE) to quickly assess where the field competition stands.
- Find "runnable code": Official code and high-star reproductions are all available—when you want to verify a paper, start here.
- Spot "inflated SOTA": Numbers on leaderboards often depend on specific data and compute—this is precisely why the home page repeatedly emphasizes "take mechanisms, not numbers."
3. Top Conference List: How the Big Five Divide Responsibilities
| Conference | Full Name | Focus | Schedule | Website |
|---|---|---|---|---|
| NeurIPS | Conference on Neural Information Processing Systems | ML comprehensive, theory, neuroscience crossover | Every December | neurips.cc |
| ICML | International Conference on Machine Learning | ML methods, optimization, probabilistic models | Every July | icml.cc |
| ICLR | International Conference on Learning Representations | Representation learning, deep learning, hot topics (attention/diffusion) | Every April–May | iclr.cc |
| CVPR | IEEE/CVF Conference on Computer Vision and Pattern Recognition | Computer Vision (images, video, multimodal) | Every June | cvpr.thecvf.com |
| ACL | Annual Meeting of the Association for Computational Linguistics | Natural Language Processing (language models, machine translation) | Every July | aclweb.org |
Top conference papers = the strongest filter of the year (typically 20%–30% acceptance rate). Practical advice: rather than daily arXiv roulette, focus on three things per year—NeurIPS/ICML/ICLR accepted papers lists (for comprehensive trends), CVPR best papers (for visual hotspots), ACL award papers (for NLP wind vane). Two hours to skim titles and abstracts each time is enough.
4. Frontier Tracking Pace
- 1–2 hours per week: scan arXiv titles → pick 10 to read abstracts → pick 2–3 to skim → pick 1 to deep-read via the three-pass method.
- Monthly review: what are the main trend lines you've seen this month? Which are bubbles, and which have real value?
- Frontier tracking tips and FAQ are at Reading Discipline & FAQ; the site's Frontier Progress has already systematically covered important breakthroughs of the 2020s (large models, diffusion models, Agents) as a starting point for tracking.
The right mindset for tracking the frontier
The value of frontier tracking isn't "knowing first"—it's "knowing more accurately than others." Being two weeks late to read a paper is fine; being a more reliable judge than someone who just retold it (how this SOTA was achieved, what it cost, where its limitations lie) is what separates advanced learners. Slow and accurate beats fast and panicked.
6. A Method for Reading a Paper: The Three-Pass Method
Regardless of which path you take, when it comes to "reading one paper," the Three-Pass Method proposed by computer scientist S. Keshav is recommended—depth increases with each pass, and each pass allows "early termination":
| Pass | Time | What to Read | Output | When to Stop at This Pass |
|---|---|---|---|---|
| Pass 1: Bird's-Eye View | 5–10 min | Title, abstract, introduction, section headings, all figures, conclusion | Answer: "What does this paper do? What's its contribution?" | It's not in your interest/direction—skip it |
| Pass 2: Close Read | ~1 hour | Methods, experiments, every detail of figures, highlight key definitions and formulas | Can retell the paper's full argument to someone else | You only need to "use" it, not "modify" it |
| Pass 3: Reconstruction | Several hours | Read line by line, reproduce derivations, question every assumption, reverse-review the experimental design | Can point out flaws and room for improvement | You're doing research, reproduction, or writing a survey |
Pass 1 5–10 min → read abstract + intro + figures + conclusion → decide: read or skip
Pass 2 ~1 hour → read methods + experiments, highlight definitions → decide: how deep to go
Pass 3 Several hours → reproduce derivations, question assumptions → decide: follow up or notFor ML papers, the three-pass method has three additional tips:
- Always read abstract → figures → conclusion first. ML paper figures have extremely high information density: one architecture diagram beats three paragraphs of text, one experiment curve chart directly exposes a method's true level. Look at figures first, then read text with questions in mind.
- Always do "ablation tracking" in Pass 2. Whenever you see "Table X: Ablation study," pause and ask: what is the author proving by removing this component? This is the core method for distinguishing "real innovation" from "engineering stacking."
- Translate "results" into "boundaries" in Pass 3. Add a qualifier to every SOTA number: under what data, what compute, what evaluation protocol does it hold? After translating, you've truly understood the paper (more methodology at Reading Discipline & FAQ).
7. Tradeoffs: Deep Read vs. Skim
The biggest resource constraint in reading papers is time, so all strategies revolve around one tradeoff: how to allocate limited energy between deep reading and skimming.
| Dimension | Skimming | Deep Reading |
|---|---|---|
| Unit cost | 10–30 min/paper | 3–8 hours/paper |
| Gain | Breadth: know "what exists, where trends are" | Depth: know "why, where boundaries lie" |
| Use case | Frontier tracking, candidate screening, building big-picture sense | Strongly related to your research direction, foundational classics, papers you'll reproduce |
| Risk | See the forest but miss the trees, echo-chamber | See the tree but miss the forest, narrowed perspective |
Practical recommendation: the 80/20 rule—handle 80% of papers with Pass 1 (bird's-eye view), 15% with Pass 2 (close read), and only 5% deserve Pass 3 (reconstruction). Papers truly worth deep reading fall into three categories:
- Paradigm-level classics: papers that changed the thinking of an entire field (e.g., Transformer, ResNet, GAN)—deep-read them, because they are the bedrock of your future understanding.
- Strongly related to your work/research: papers that directly address your current problem—deep-read them, study their ablations and limitations, judge whether you can adapt them to your scenario.
- Repeatedly cited but controversial: since you can't avoid them and there's disagreement—deep-read them, make your own judgment.
The rest (hot news, marginal improvements, pure engineering reports) should all be skimmed. Deep reading is investment; skimming is expense—using expense to maintain breadth and investment to gain depth is the most sustainable reading strategy in this field.
How to coordinate with other sections on this site
While reading paths, treat the site's four companion articles as an "accelerator": the Paper Map helps you judge a paper's position on the map (source / milestone / survey), deciding whether to skim or deep-read; Classic Paper Deep Dives does Pass-3-level breakdowns of the most important papers for you; Frontier Progress helps you place new papers on trend lines; Reading Discipline & FAQ handles "can't understand, can't remember, can't persist."
8. Further Reading
At this point, you have three paths and the three-pass method. Next, choose your next destination based on your goal:
- Start Here — Main entry point for the papers section: three values of reading papers and one crucial piece of advice
- Paper Map — Macro coordinate system for systematic study: 60 years of key papers organized by theme and timeline
- Classic Paper Deep Dives — Step-by-step breakdowns of the classic papers in the Quick Start and Systematic Study paths
- Frontier Progress — Starting point for frontier tracking: systematic overview of important 2020s breakthroughs
- Reading Discipline & FAQ — Supplement to the three-pass method: note-taking, what to do when you can't understand, how to judge paper quality
- Evolutionary History — Put papers into a bigger timeline: the historical narrative thread of machine learning
- Learning Paths — Return from papers to the system: the site's complete knowledge learning routes
- Resources — Real resource collection for paper code reproduction, datasets, and tools
References
All materials below are real, publicly accessible resources for in-depth self-study:
- arXiv Preprint Library — Primary publishing venue for most ML papers, browse by
cs.LG/cs.CV/cs.CLcategories - Papers with Code — Paper + code + SOTA leaderboard aggregation, primary battlefield for frontier tracking and reproduction
- Google Scholar — Citation tracking (Cited by), follow the timeline to find follow-up work
- NeurIPS website — Comprehensive ML top conference, every December
- ICML website — ML methods top conference, every July
- ICLR website — Representation learning and deep learning top conference, every April–May
- CVPR website — Computer Vision top conference, every June
- ACL website — NLP top conference, every July
- Keshav. How to Read a Paper (2007) — The original paper of the "three-pass method," a must-read for reading methodology
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — Original Transformer paper
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019)
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020) — GPT-3
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (ICLR 2023)
- Krizhevsky et al. ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012) — AlexNet
- He et al. Deep Residual Learning for Image Recognition (CVPR 2016) — ResNet
- Goodfellow et al. Generative Adversarial Nets (NeurIPS 2014) — GAN
- Ho et al. Denoising Diffusion Probabilistic Models (NeurIPS 2020) — DDPM, a representative diffusion model paper
- Silver et al. Mastering the game of Go with deep neural networks and tree search (Nature 2016) — AlphaGo
- Jay Alammar. The Illustrated Transformer (2018) — A classic example of visual deep-dive reading of the Transformer paper