Theme
Learning Paths: Three Routes
Straight to the point: this site has over 50 pages. It is a library you can pull from as needed, and a map you can unfold in order — but it is definitely not a book you read cover to cover. Reading "from page one to page fifty" with no target in mind will leave you forgetting most of it after two weeks, and the content density will discourage you. The right way to use it: pick a route, read only what that route requires, and save the rest for a future you. If you're not yet sure whether deep learning is worth investing in, start with What is Deep Learning; if you want to see the full six-part structure of this project at a glance, see Deep Learning Architecture Anatomy.
All pages on the site fall into four categories by content type. Understanding these four tells you what mindset to bring when reading any given page:
- Building concepts: the guide and concepts sections answer "what is this and why" — these need close reading;
- Explaining mechanisms: the case-studies section answers "how does this architecture work on a real task" — these need close reading plus hands-on reproduction;
- Providing evidence: the papers section answers "what exactly does the paper say and what's the evidence" — these need critical reading;
- Changing behavior: the practice, career, and resources sections answer "what should I do, say, and look up" — these call for action, not memorization.
Below are three routes. First, a quick self-assessment: if your goal is to land an algorithm role within a few weeks, take Route One. If your goal is to build a solid foundation so you can independently read papers and reproduce projects, take Route Two. If you're already working on projects but frequently get stuck, jump straight to Route Three's index table.
Route One: Job-Hunting Sprint (~2 Weeks)
Who it's for: People with Python and foundational machine learning background, aiming to reach a state of "can explain principles, can write demos, can handle interviews" within 2–3 weeks. The career module serves as the backbone, but the first 8 days focus on filling in conceptual foundations — interviewers won't ask fewer principle questions just because you're sprinting. Recommended daily investment: 4–6 hours.
| Day | Content | Acceptance Criteria for the Day |
|---|---|---|
| Day 1–2 | What is Deep Learning, Neural Network Fundamentals | Can explain "what deep learning is doing" to a non-technical friend; can hand-draw the computational graph for single-layer forward propagation |
| Day 3–4 | Backpropagation and Automatic Differentiation, Optimization and Gradient Descent | Can derive backpropagation for a two-layer network by hand; can explain the differences among SGD / Momentum / Adam |
| Day 5 | Loss Functions and Output Layers, Overfitting and Regularization | Can name loss choices for classification vs. regression; can list 4+ regularization techniques |
| Day 6–7 | CNN and Computer Vision, Transformer Architecture overview | Can compute the parameter count for a convolution output; can explain what Q, K, V are in self-attention |
| Day 8–10 | Complete an end-to-end demo (Building a Deep Learning Project from Scratch or Progressive Tutorial: Three Versions That Work) | Train and evaluate a model independently, push code to GitHub |
| Day 11–12 | JD Knowledge Breakdown, Competency Benchmarking: What to Highlight on Your Resume | Write down 5 strengths and 3 gaps against the JD Checklist |
| Day 13–14 | Interview Question Bank mock practice + fill in knowledge gaps | Can answer 80% of high-frequency questions from memory without notes; record a self-introduction and review to improve |
Three acceptance criteria at the end of two weeks: one showcaseable project (can explain data, model, loss, and evaluation); roughly 70% "conceptual-level mastery" of the JD checklist; and a resume rewritten for your target role. See the career module primer in the further reading section for the full landscape of roles and complete guides.
A Note for Sprinters
The sprint route sacrifices depth — the trade-off is having "breadth" without "roots." If an interviewer asks "why is it designed this way," probing two or three levels deep is enough; you don't need to dive into every page's details — but you should at least read through Training Recipes and Hyperparameter Tuning, because "what to do when loss doesn't decrease" is the most common interview scenario.
Route Two: Systematic Deep Dive (~8 Weeks)
Who it's for: People who want to truly internalize deep learning, with 10–15 hours per week to invest. This route unfolds in dependency order — each week builds on the previous ones, so don't skip weeks. If time is tight, stretch the timeline rather than compressing weekly content.
| Week | Topic | Content and Entry Points | Acceptance Criteria for the Week |
|---|---|---|---|
| Week 1 | Math Foundations | Math Primer, Neural Network Fundamentals, Backpropagation and Automatic Differentiation | Independently derive gradients for a three-layer network by hand; understand vector/matrix differentiation rules |
| Week 2 | Training Recipes | Loss Functions and Output Layers, Initialization and Normalization, Overfitting and Regularization, plus training recipes and tuning | Can name 5 troubleshooting directions for "loss won't decrease"; can explain the difference between BN and LN |
| Week 3 | Data and Evaluation | Data and Data Engineering, Deep Learning Evaluation and Experiments, Evaluation in Practice | Can design a data strategy and evaluation metrics for a given task; can identify data leakage |
| Week 4 | Vision and Sequences | RNN and Sequence Modeling, Transformer Architecture (paired with CNN content to understand inductive bias differences) | Can compare the trade-offs among RNN, CNN, and Transformer architectures |
| Week 5 | Generation and Representation | VAE and GAN, Diffusion Models and Generative AI, Representation Learning and Pretraining, Generative Models | Can explain the principle differences among the three generative model types; can describe the "pretrain → fine-tune" paradigm |
| Week 6 | Frontier Case Studies | Large Language Models (LLMs), Multimodal Models, plus optional picks like Graph Neural Networks, Speech and Audio based on interest | Can draw the flowchart of LLM pretraining through alignment (fine-tuning / RLHF) |
| Week 7 | Paper Deep Dives | First read Reading Paths and Paper Map, then deeply read 2–3 papers from Classic Paper Deep Dives | Can write four-column notes ("Problem – Method – Evidence – Limitations") for each paper |
| Week 8 | Project Integration | Portfolio Projects, MLOps and Model Deployment | Complete a deployable end-to-end project with a well-written README and experiment logs |
At the end of eight weeks, you should possess three abilities: independently read most papers published before 2018; build a reasonable baseline from scratch and systematically tune it; make stable judgments about "which model fits which scenario." Keep your skills sharp by making Frontier Advances and Common Pitfalls and Anti-Patterns part of your weekly reading routine.
Route Three: Desk Reference
Who it's for: People already working on projects but unsure where to look up answers to specific problems. This is a "problem → page" index table — bookmark this page and come back anytime.
| The Problem You're Facing | Go Directly Here |
|---|---|
| Loss won't decrease, training diverges, gradient explosion/vanishing | Debugging and Diagnosis, Optimization and Gradient Descent |
| Overfitting: validation set much worse than training set | Training Recipes and Hyperparameter Tuning |
| Don't know which architecture to choose for images | CNN and Computer Vision |
| Working with text, sequences, long context | Attention Mechanism |
| Need to fine-tune or call a large model | Large Language Models (LLMs) |
| Deploying a model, putting it on cloud, monitoring | MLOps and Model Deployment |
| Data is dirty, too few samples, class imbalance | Data and Data Engineering, Datasets and Tool Archives |
| Can't understand formulas or math notation | Math Primer, Glossary |
| Results look good but can't explain why | Interpretability and Fairness |
| Preparing for interviews, job hunting | Interview Question Bank, JD Knowledge Breakdown |
| Want to build a portfolio-worthy project | Portfolio Projects, Building a Deep Learning Project from Scratch |
| Don't know how to read papers | Reading Paths, Reading Discipline and FAQ |
| Framework selection (PyTorch / JAX / inference engines) | Framework and Tool Comparison |
| How to scientifically evaluate model performance | Evaluation in Practice, Deep Learning Evaluation and Experiments |
This table covers the 14 most frequently asked problem categories. If your problem isn't here, there are two fallback entries: DL Design Principles (deduce what to do from first principles) and Common Pitfalls and Anti-Patterns (a list of pitfalls others have already fallen into).
Timeliness Note
The site's "freshness" strategy: conceptual, historical, and theoretical pages are largely evergreen. However, sections that involve tech stacks, model scores, roles, and tools (such as the JD Checklist, Curated Resource List, and Frontier Advances) will display a dataAsOf timestamp in the frontmatter — trust the date shown on each page. The time estimates for the three routes are based on the median daily investment — you can linearly scale them up or down to match your own pace. Route One compressed to 10 days or stretched to 30 days both work; the key is to hold onto acceptance criteria, not to fixate on the number of days.
One more note: any roadmap is only a "minimum viable path." Real competitiveness comes from two things outside any roadmap — getting your hands dirty with code (Progressive Tutorial: Three Versions That Work provides a shallow-to-deep runway) and reading primary papers (Week 7 of Route Two is just the starting point). When your questions start exceeding the boundaries of this map, it means you've entered the stage of "drawing your own maps" — and that's when real learning begins.
Further Reading
- Deep Learning Architecture Anatomy — understand the full picture of "one problem → one complete pipeline"
- DL vs ML vs AI vs Traditional Methods — clarify conceptual boundaries before picking a route
- A Brief History of Deep Learning — place each topic within a timeline for deeper understanding
- Paper Map — the "treasure map" for the second half of the systematic deep dive
- Glossary — a dictionary you can consult anytime along any route
- career module primer — the entry point for the job-hunting sprint route
References
This page is a site navigation page with no independent academic citations; all content referenced in the routes is linked above. Two additional publicly available learning resources worth following long-term:
- Goodfellow, Bengio, Courville. Deep Learning (MIT Press 2016) — an optional supplementary textbook for the systematic deep dive
- fast.ai. Practical Deep Learning for Coders — a mature reference for the "hands-on first, theory second" route