Theme
Learning Paths: Three Routes
This site has 50+ pages. If you read from the first page straight through, you'll most likely zone out by the fourth article in the core knowledge section—not because the content is bad, but because reading cover to cover is not the correct approach to learning machine learning as a discipline. You don't need a library; you need a roadmap.
The content on this site actually has only four types: concept-building (reading guides), mechanism-explaining (core knowledge), evidence-providing (cases and papers), and behavior-changing (practice and career). People in different situations should consume them in completely different orders and depths. This page lays out three pre-assembled routes—pick the one that matches your situation.
text
What's your situation?
│
┌───────────────┼───────────────────┐
▼ ▼ ▼
Interview within Plenty of time, Already doing ML
1-3 months want to transition
│ │ │
┌───────────────┐ ┌───────────────┐ ┌──────────────┐
│ ① Career Sprint│ │ ② Systematic │ │ ③ Desk Ref. │
│ ~2 weeks │ │ Track │ │ Track │
│ │ │ ~8 weeks │ │ Use on need │
├───────────────┤ ├───────────────┤ ├──────────────┤
│ career main │ │ Guide→Knowledge│ │ Encountering │
│ +high-freq │ │ →Cases→Practice│ │ a problem? │
│ concepts │ │ →Papers │ │ →Check index │
│ +hands-on demo│ │ Weekly checks │ │ →Back to work│
├───────────────┤ ├───────────────┤ ├──────────────┤
│ Deliverable: │ │ Deliverable: │ │ Deliverable: │
│ revised resume│ │ Can explain │ │ Current │
│ + 10 practice │ │ the full body of knowledge + │ │ problem solved│
│ questions │ │ portfolio │ │ │
└───────────────┘ └───────────────┘ └──────────────┘The three routes are not mutually exclusive
Many people's actual trajectory is: use the desk reference for a few months, then run through the systematic track after deciding to transition, then run through the second half of the sprint track right before an interview. Routes are navigation, not rails.
Route 1: Career Sprint Track (~2 weeks)
Target audience: People with machine learning/algorithm interviews within 1–3 months, who already have basic programming skills (Python), don't have large blocks of time, and only have evenings and weekends available.
Core idea: Interviews fundamentally test two things—can you explain the machine learning framework clearly (from probability to models to engineering), and can your resume and communication pass the screening? So this route works backward from the end goal, deriving what to study from job descriptions and interview questions. It doesn't aim for completeness—it aims to cover all high-frequency test points and be able to explain "why" for each one.
Main thread: four career module articles
| Order | Page | What you should take away |
|---|---|---|
| 1 | JD List: Open Positions at Major Companies | What the target roles actually require—note the skill keywords that appear repeatedly |
| 2 | Skills Map | Map the JD skill keywords to pages on this site, generating your personal study checklist |
| 3 | Resume Analysis | Rewrite your project experience following the standards here, highlighting models and business decisions |
| 4 | Interview Questions | Use this last for self-testing, not on day one |
Side dishes: most frequently tested concepts + one hands-on demo
Based on the distribution of JDs and interview questions, the following four articles cover the vast majority of "algorithm-side" test points, read them in this order:
- Model Evaluation and Validation—almost universally tested in every ML interview; overfitting, bias-variance, AUC, and cross-validation should come out naturally;
- Supervised Learning—principles and derivations of linear/logistic regression, SVM, decision trees;
- Overfitting and Regularization—differences and use cases for L1/L2, Dropout, and early stopping;
- Optimization and Gradient Descent—gradient vanishing/exploding, learning rate tradeoffs, SGD vs. Adam.
If you can only add one more, add Deep Learning Fundamentals—in the era of large models, nearly every role will ask about it.
A resume with no hands-on projects lacks credibility. Use your weekend to run through Building an ML Project from Scratch: the v1 minimal linear regression in the repo is only a few dozen lines, v2 adds feature engineering and evaluation, v3 switches to gradient boosting trees. Run through each layer and tweak them, and you'll have a personal project you can talk about for five minutes nonstop in an interview.
Two-week rhythm
Week 1 (understanding + resume):
- Monday to Wednesday: JD list → skills map, produce your study checklist;
- Thursday to Friday: read the four concept articles in order above; after each one, close it and use your phone to record yourself explaining the core mechanism. Re-read sections where you stumble;
- Weekend: run build-your-own from v1 to v3; revise your resume draft according to the standards on the resume analysis page.
Week 2 (communication + self-testing):
- Monday to Wednesday: interview questions, answer each one yourself before checking the answer; mark anything you couldn't answer and revisit the corresponding concept page;
- Thursday to Friday: ask a friend to do mock interviews or do self-Q&A with voice recording, focusing on "why" questions (why does L1 produce sparse solutions? why don't tree models need normalization?);
- Weekend: finalize your resume, organize the demo project into a repo that can be demonstrated.
Acceptance criteria after two weeks
You should have three things in hand: a resume rewritten to the standards of an algorithm role, a running and explainable ML project repo, and 10 interview questions with answers that articulate tradeoffs and derivations. If you're missing one, it means you cut corners in the corresponding step.
Route 2: Systematic Track (~8 weeks)
Target audience: People with relatively ample time (can invest 8–10 hours per week), who want to transition or build a solid foundation within 1–2 quarters, and value depth over speed.
Core idea: Machine learning knowledge has a dependency graph—without understanding evaluation, you can't understand regularization; without understanding feature engineering, you can't appreciate the nuance in cases; without dissecting cases, you can't write your own design principles. This route is laid out in dependency order, with weekly checkpoints—you shouldn't advance to the next week until you pass the current one.
| Week | Content | Checkpoint |
|---|---|---|
| Week 1 | Five reading guides: What Is ML → ML vs AI → History → Anatomy → Paths | You can explain to someone who's never studied ML "the difference between machine learning and traditional programming," and draw the four steps of an ML project (data → model → algorithm → evaluation) |
| Week 2 | Math foundation: linear algebra and probability sections of Math Primer + Supervised Learning | You can derive the least-squares solution for linear regression by hand; explain the meaning of maximum likelihood estimation |
| Week 3 | Evaluation and regularization: Model Evaluation and Validation → Overfitting and Regularization → Optimization and Gradient Descent | You can explain the bias-variance tradeoff and name at least three ways to handle overfitting; articulate the differences between gradient descent, SGD, and Adam |
| Week 4 | Features and data: Feature Engineering → Data and Data Engineering; optional cases: Tree Models and Ensemble Learning, Linear Models | Given any tabular dataset, you can list a complete feature engineering checklist (missing values, encoding, scaling, interactions) |
| Week 5 | Unsupervised: Unsupervised Learning + case Clustering and Dimensionality Reduction; hands-on: Building an ML Project from Scratch v1→v3 | You can explain the principles and limitations of K-Means and PCA; the demo runs and you can explain what each step does |
| Week 6 | Deep learning: Deep Learning Fundamentals → cases CNN, Transformer | You can draw forward/backward propagation for a multi-layer perceptron by hand, and explain the core ideas of convolution and attention |
| Week 7 | Practice wrap-up: Hyperparameter Tuning → Design Principles → Common Pitfalls → Evaluation in Practice | Your project can run on someone else's machine, and you can explain what kind of inputs it would fail on |
| Week 8 | Papers + career materials wrap-up: deep-read at least 3 papers from Classic Papers (e.g., AlexNet, Attention, ResNet), skim Frontier; then go through the career module | You can explain the motivation and contributions of at least two classic papers; your resume and 10 question answers meet Route 1's acceptance criteria |
The hard requirement for Week 5
The only week that cannot be compressed in the eight weeks is the hands-on week. After reading concept articles, you'll get the illusion that "I already understand this"—that illusion shatters when you face real dirty data for the first time. Without having been personally schooled by NaN, class imbalance, and overfitting, your transition isn't complete.
If time is truly tight, you can cut the deep learning in Week 6 and the papers in Week 8 in half, but don't cut the reading guide, evaluation, and hands-on sections—they are the load-bearing walls of this route.
Route 3: Desk Reference Track (for working practitioners)
Target audience: People already working in ML-related development, where problems show up in units of "I'm stuck today," with no need or patience for reading cover to cover.
The way to use this route: when you encounter a problem, look up the entry page from the table below, read it to solve the problem, and move on—don't go down rabbit holes. Save rabbit holes for the weekend.
| Problem you're facing | Read this |
|---|---|
| Model does great on training set, terrible on test set | Overfitting and Regularization, Model Evaluation and Validation |
| Don't know what metric to use for model evaluation (classification/regression/ranking) | Model Evaluation and Validation |
| Data has many missing values and categorical variables, don't know how to handle | Feature Engineering, Data and Data Engineering |
| Training is slow, doesn't converge, loss oscillates | Optimization and Gradient Descent, Hyperparameter Tuning |
| Class imbalance, positive samples are extremely rare | Model Evaluation and Validation (class imbalance section) |
| Should I use XGBoost or a neural network | Tree Models and Ensemble Learning, How to Choose Frameworks and Tools |
| Want to quickly run an end-to-end project | Building an ML Project from Scratch, Progressive Tutorial |
| Model performance degraded after deployment | MLOps and Model Deployment (monitoring and drift detection) |
| Boss/client asks "why did the model make this decision?" | Interpretability and Fairness |
| Don't know where to start with deep learning models | Deep Learning Fundamentals, CNN, Transformer |
| Want to understand how large models are trained and fine-tuned | Large Language Models (LLM) |
| Interviewing or being interviewed, need to assess framework knowledge | Interview Questions |
| Stuck on a term | Glossary |
| Can't understand the math formulas | Math Primer |
| Looking for datasets and tools | Datasets and Tool Archives, Curated Resource List |
Bookmark this page
The value of the desk reference is its reuse rate. We recommend bookmarking this page (rather than any concept article)—it's the switchboard for the entire site.
Two notes
Paths are not rigid rules. The weeks and order for all three routes are estimated for typical backgrounds: someone with a math background can skip the math foundation, someone with strong engineering skills can fast-forward through data engineering, and a product manager transitioning into the field should spend double time on the reading guides and cases. The criterion is always whether you can pass the weekly checkpoints, not which page the calendar has turned to.
Beware of content staleness. Machine learning is a field with a short half-life—leaderboards, model scores, and tool choices all become outdated. Some pages in the career module and case sections are marked with dataAsOf (data cutoff month) in the frontmatter and also displayed at the top of the page; before citing specific roles, scores, or product features from them, check that date first. Treat data older than a quarter as "trend reference" rather than "current fact," and verify against official sources. Conceptual content (what is supervised learning, the bias-variance tradeoff) degrades much more slowly and can be relied on long-term.
Further reading
- What Is Machine Learning—the common starting point for all three routes
- Anatomy of an ML System—full-site content map, cross-referencing with this article
- Career and JD Analysis—entry point for Route 1's main thread
- Classic Papers Deep Dive—original papers for Week 8 of Route 2
- Design Principles—a page worth revisiting and cross-referencing after completing any route