Theme
Awesome Resources
Machine learning is a discipline you can't learn by just "reading", but relying entirely on your own trial and error — debugging, hair-pulling, and hyperparameter tuning — is equally terribly inefficient. Over the past decade, this field has accumulated a set of exceptionally high-quality, mostly free official resources — and they have clear difficulty progression and stylistic divisions among themselves. This page organizes them into eight categories by learning phase, each with a one-sentence positioning, who it's for, and an official link, all individually verified.
How to use
This list isn't meant for you to bookmark and forget — it's the "supplementary resource layer" paired with the site's Learning Paths. The recommended approach: first build your conceptual framework on the site (start from What Is Machine Learning and the Glossary), then pick one or two main threads to master according to this page's rhythm — trying to learn everything at once means mastering nothing; finishing one main thread from start to end beats skimming the first three chapters of ten different resources.
I. Introductory Courses
The goal at the introductory phase isn't to "learn every algorithm", but to build intuition + run code with your own hands: know what overfitting is, what training/test sets are, and why evaluation metrics matter. Three mainstream routes offer complementary styles — pick one as your main thread.
Andrew Ng's Machine Learning Specialization (Coursera)
One-sentence positioning: The world's most popular introductory ML course, a classic from 2012 that was rebuilt in 2022.
Andrew Ng's Machine Learning Specialization consists of three courses: supervised learning (linear/logistic regression, neural networks, decision trees), advanced learning algorithms and ML best practices, and unsupervised learning and recommendation systems/reinforcement learning. Its greatest strength is maximizing "intuition" — before every formula appears, it first uses visualizations and real-life examples to help you understand "why it exists". Its weakness is limited depth; the course barely covers content relevant to the current large-model era, but that's precisely its positioning as a cornerstone.
- Who it's for: Zero-basis beginners or anyone who only knows Python; interviewees who need "what ML actually does" explained clearly.
- Official link: coursera.org/specializations/machine-learning-introduction (financial aid available, or audit for free)
fast.ai's Practical Deep Learning for Coders
One-sentence positioning: Opposite to Andrew Ng's "theory first, then code", it "gives you a working model first, then opens it up so you can see what's inside".
Jeremy Howard's free course uses a top-down teaching approach: by the second lesson, you're already training and deploying an image classifier. Each subsequent lesson digs one layer deeper, eventually covering CNNs, NLP, tabular data, collaborative filtering, random forests, and even Stable Diffusion. The companion book Deep Learning for Coders with fastai and PyTorch is freely readable online.
- Who it's for: People with programming experience (Python is enough) who want to build things quickly; hands-on learners who were intimidated by math.
- Official link: course.fast.ai
Li Mu's Dive into Deep Learning (D2L)
One-sentence positioning: The only deep learning textbook that achieves "every formula paired with runnable code", available in both Chinese and English, used by 500+ universities worldwide.
D2L takes "hands-on" to the extreme: the entire book is composed of Jupyter Notebooks, from implementing linear regression from scratch to Transformers and BERT — every section goes "derive the formula → write code → see results". The Chinese version is maintained by Li Mu's team, with Bilibili companion videos in full Chinese — the top learning path for deep learning in China.
- Who it's for: Those with basic programming skills who want to seriously learn deep learning (not just tabular ML); people who want to build intuition through code reproduction.
- Official link: zh.d2l.ai (Chinese version) | github.com/d2l-ai/d2l-zh
How to choose among the three
For interviews and wanting a systematic conceptual understanding → Andrew Ng; for building the fastest demo-able project → fast.ai; committed to going deep into deep learning / large models → go straight to D2L. The three are not mutually exclusive — many people start with Andrew Ng and deepen with D2L. At the introductory phase, pair it with the site's Math Primer for on-demand gap-filling rather than spending a month studying math first.
II. Classic Textbooks
A textbook's value lies not in "newness", but in complete systems and rigorous derivations — they answer "why" rather than "how to tune". The five below are ordered from easy to hard, all with official free access channels.
An Introduction to Statistical Learning (ISL)
One-sentence positioning: The most approachable machine learning textbook from a statistical perspective, with low math barriers throughout and a focus on application.
ISL is co-authored by Stanford's Hastie, Tibshirani, and four other statisticians, stringing together regression, classification, resampling, regularization, tree models, SVMs, and deep learning with minimal formulas. The second edition (2021, R version) and Python version (2023) are both published, with free PDF on the official website, and each chapter comes with complete exercise code — this is its biggest differentiator from ESL below: it can serve as a hands-on textbook.
- Who it's for: Beginners who want to understand the "statistical meaning" of models; people working with tabular data, risk control, and marketing analytics.
- Official link: statlearning.com (free PDF + Python/R resources)
The Elements of Statistical Learning (ESL)
One-sentence positioning: The advanced sequel to ISL, the "Bible" of statistical learning, with an order of magnitude higher formula density.
The same authors' companion volume. If ISL is "helping you use it", ESL is "helping you understand why doing it this way is correct". It covers nearly all classic content: supervised/unsupervised, kernel methods, trees and boosting, graphical models, random forests, and more — it is the citation source for countless papers. The authors provide a free PDF on the official website.
- Who it's for: Advanced learners with some math foundation (linear algebra + probability) who want to dig deep into algorithm principles; people preparing to read papers.
- Official link: hastie.su.domains/ElemStatLearn
Deep Learning ("The Deep Learning Book" / "Flower Book")
One-sentence positioning: The foundational textbook of deep learning, co-authored by the three giants Goodfellow, Bengio, and Courville, freely readable online.
"The Flower Book" (named after its cover) fills in linear algebra, probability, numerical computation, and ML fundamentals in its first 5 chapters — this part remains, to this day, the best mathematical prerequisite for deep learning. The second half covers feedforward networks, regularization, optimization, CNNs, RNNs, representation learning, and generative models. The full book is freely readable on the official website; a Chinese translation is available from Posts & Telecom Press.
- Who it's for: Those learning deep learning who want to systematically fill in math and theory; people seeking answers to questions like "why does regularization work" and "why is optimization hard".
- Official link: deeplearningbook.org
Pattern Recognition and Machine Learning (PRML)
One-sentence positioning: An encyclopedia of machine learning from a Bayesian perspective, unifying all of traditional ML through the skeleton of probabilistic graphical models.
Christopher Bishop's PRML takes another "worldview": it puts every model into the Bayesian framework (prior → likelihood → posterior) and systematically introduces probabilistic graphical models, variational inference, and expectation propagation. The elegance of its formulas makes countless people treat it as their "red book", and it is an essential path for understanding the intellectual roots of modern generative models. The author offers a free PDF on his Microsoft Research page.
- Who it's for: Advanced learners interested in probabilistic/statistical methods who can handle formulas; those preparing for research-oriented job interviews who need systematic theory review.
- Official link: microsoft.com/en-us/research/publication/pattern-recognition-machine-learning
Reinforcement Learning: An Introduction (Sutton & Barto)
One-sentence positioning: The most cited textbook in the reinforcement learning field; the second edition (2018) is offered as a free PDF.
Sutton and Barto are the founding fathers of RL. This book starts from multi-armed bandits, progresses through dynamic programming, Monte Carlo, temporal difference, Q-learning, policy gradients, function approximation, and on to methods of the AlphaGo era. It explains core challenges like "exploration-exploitation" and "credit assignment" extraordinarily thoroughly. The full second edition is freely downloadable from the authors' website.
- Who it's for: Anyone who wants to engage with RL (read this before touching Deep Q-Network and PPO implementations); people who want to understand the underlying logic of large-model RLHF.
- Official link: incompleteideas.net/book/the-book-2nd.html (free PDF)
Textbook Pairing Advice
Entry → ISL (hands-on) + Flower Book chapters 1–5 (fill in math); Advanced → ESL or PRML (pick one, different styles: ESL leans statistical, PRML leans Bayesian); Specialized → RL: Sutton & Barto. Do not attempt to read all five cover to cover — they are reference books and reference frames, not a reading marathon.
III. Deep Learning Specialized Courses
Introductory courses solve "how to train deep learning"; these courses solve "what exactly is a given direction". They are all public courses from Stanford or top researchers, with content, assignments, and slides all freely available.
CS231n: Stanford's Deep Learning for Computer Vision
One-sentence positioning: The required course for the computer vision direction, and also "the course that explains CNNs most clearly".
Starting from image classification, it systematically covers every detail of neural network training (activation functions, BatchNorm, regularization, optimizers, transfer learning), then goes deep into CNN architecture evolution (AlexNet → VGG → ResNet → object detection/segmentation). Assignments require implementing backpropagation from scratch in numpy and writing a two-layer convolutional network by hand. The course notes (cs231n.github.io) themselves serve as a deeply bookmarked deep learning intro document.
- Who it's for: Those who have already run some deep learning code and want to go deep into vision, or those who want to thoroughly understand backpropagation details.
- Official link: cs231n.stanford.edu
CS224n: Stanford's Deep Learning for Natural Language Processing
One-sentence positioning: The complete NLP course from word embeddings all the way to Transformers and LLM alignment.
Content covers word2vec, RNN/LSTM, attention, Transformer, BERT pretraining, fine-tuning, RLHF/DPO alignment, RAG and agents. After the mid-2020s, the course has been fully updated to center on LLMs, with the final project being to hand-write a minimal GPT. Lecture notes and slides are updated in real time and publicly available, with complete recordings from past years on YouTube.
- Who it's for: Those entering the NLP / LLM application space; those who want to understand the complete "pretraining → fine-tuning → alignment" chain.
- Official link: web.stanford.edu/class/cs224n
Karpathy's Neural Networks: Zero to Hero
One-sentence positioning: The video series across the internet that explains "exactly how neural networks compute internally" the most thoroughly — handwritten code from scratch, building a GPT all the way.
Andrej Karpathy (OpenAI co-founder) takes you through a video series that starts from implementing micrograd (automatic differentiation) and makemore (character-level language model), then writes from attention all the way to a complete GPT and Tokenizer. No slide deck is more intuitive than "typing code on screen, explaining backpropagation line by line". Completely free, publicly available on YouTube.
- Who it's for: Those using neural networks as a black box but feeling uneasy about it; those who want to understand large models from the bottom up (not just call APIs).
- Official link: karpathy.ai/zero-to-hero.html
IV. Practice & Competitions
Reading and courses can't provide two things: the messiness of real data and the pressure of rankings. These two gaps are respectively filled by Kaggle and Hugging Face.
Kaggle: Competitions, Courses & Datasets — One Platform
One-sentence positioning: The world's largest machine learning competition community — "compete with tens of thousands of people worldwide on real data".
Kaggle's three core assets: Competitions (real business problems + public leaderboards, the fastest path to practice and portfolio building), Learn (short courses covering Python, Pandas, ML, deep learning, SQL — 1–2 hours each), and Datasets (a massive open dataset repository). Start by "reading a gold-medal solution" — publicly shared notebooks and discussions are the best teaching material.
- Who it's for: Anyone who wants to verify "whether what I've learned actually works"; students needing portfolio pieces for job hunting.
- Official link: kaggle.com/competitions | kaggle.com/learn | kaggle.com/datasets
Hugging Face Courses: From Transformers to Agents
One-sentence positioning: The official teaching system for the open-source AI ecosystem (Transformers, diffusers, agent frameworks).
Hugging Face offers multiple free courses: NLP Course (from fine-tuning BERT to training your own model, entirely based on the Transformers library), image generation/audio courses, and a course for Agent development. Its difference from Kaggle: Kaggle practices "competition playing", HF practices "building products with modern toolchains".
- Who it's for: People who need to use/fine-tune open-source large models; those doing LLM application development (RAG, Agents).
- Official link: huggingface.co/learn/nlp-course
The Right Approach to Practice
When doing Kaggle, focus on three metrics: understand one gold-medal solution → reproduce it → improve on it — this is more effective than blindly entering ten competitions. The commercial value ranking of competitions is: understanding solutions on the leaderboard > submitting a few times > winning an award. Also recommended: read the site's Framework Comparison before deciding your main framework, to avoid repeatedly jumping between tool choices.
V. Framework & Library Official Documentation
Framework docs are the authoritative source for "look up as you go", but knowing how to read docs and being drowned by docs are two different things. Here we keep only five official doc entry points you must encounter when learning ML.
scikit-learn: The De Facto Standard for Classic ML
One-sentence positioning: Python's most comprehensive classic ML library, one API spanning data preprocessing, models, evaluation, and pipelines.
scikit-learn's documentation is one of its greatest strengths: each algorithm has a three-part structure of "plain explanation + mathematical definition + complete example". Its User Guide is essentially a "practical ML handbook", and the example library (auto_examples) covers virtually all common scenarios. 90% of the workflow for tabular data and small-to-medium tasks can be done within sklearn.
- Official link: scikit-learn.org/stable
PyTorch: The Mainstream Deep Learning Framework
One-sentence positioning: The current most mainstream deep learning framework in both academia and industry: the trio of "tensors + autograd + modular design".
PyTorch's official docs are divided into three parts: Tutorials (from a 60-minute intro to Transformer implementations — the first stop for beginners), Docs/API (module-level authoritative reference), and PyTorch Ecosystem (surrounding library index). 90% of deep learning pitfalls (tensor shapes, device transfer, gradient accumulation) have counterparts in the doc examples.
- Official link: pytorch.org
XGBoost & LightGBM: The Dual Champions of Gradient Boosting Trees
One-sentence positioning: The "de facto kings" of tabular data tasks, perennial champion models in Kaggle structured competitions.
XGBoost docs are famous for their tutorials — especially An Introduction to Boosted Trees, which explains boosting's math accessibly to an almost astonishing degree. LightGBM docs are equally concise, with its histogram algorithm, Leaf-wise growth, and native categorical feature support being frequently consulted points for engineering tuning. The two libraries have highly similar APIs (both provide sklearn-style interfaces) — knowing one means knowing both.
- Official link: xgboost.readthedocs.io | lightgbm.readthedocs.io
MLflow: Engineering Foundation for Experiment Management and Model Lifecycles
One-sentence positioning: The engineering tool that turns "models in notebooks" into formal, reproducible, trackable, deployable products.
MLflow docs are organized around four main threads: Tracking (recording hyperparameters/metrics/artifacts of each experiment), Models (unified model packaging format), Model Registry (model versioning and approval for deployment), and Projects (reproducible training environments). When your experiment count exceeds double digits, or when you need to collaborate with a team, MLflow is the answer — it solves engineering problems like "turning models from notebooks into production services".
- Official link: mlflow.org/docs
How to Read Docs Without Wasting Time
Framework docs are not textbooks — don't read them cover to cover. The right approach: spend 30 minutes reading Tutorials to build a mental model, then go into the API pages with specific questions ("what are this function's parameters"), and finally use the example pages as templates to modify your own code. Framework tradeoffs can be found in the Framework Comparison.
VI. Dataset Resources
A model's ceiling is determined by its data. During practice and learning, prioritize "clean, small, paper-backed" classic datasets; find domain-specific real data when doing projects.
The site has a standalone Dataset Guide covering commonly used classic datasets and retrieval methods; here we give only four official channels for acquiring data:
- Kaggle Datasets: largest volume, best search experience among open dataset platforms, competition data mostly published here, many with runnable notebook baselines.
- UCI Machine Learning Repository: the oldest classic dataset warehouse (Iris, Adult, Wine Quality, etc. all come from here), great for teaching and paper reproduction, with standardized formats and authoritative sources.
- Hugging Face Datasets: the default dataset platform of the NLP / multi-modal era, seamlessly integrated with the Transformers ecosystem — you can almost find every mainstream paper's official data version here.
- Papers with Code Datasets: dataset catalog organized by paper, each dataset listing "SOTA papers that cite it" — the first choice for finding benchmark data.
Two Reminders About Datasets
First, copyright and licensing: read the license before using new datasets, especially user-generated content obtained by scraping; second, data leakage: cross-file merging, global statistics, and future information can all artificially inflate offline metrics — these pitfalls are discussed in detail in the evaluation, feature engineering, and engineering practice sections, or you can look them up by term in the Glossary. Beginners should start with classic datasets like Iris, Titanic, and house price prediction.
VII. Communities & News Stays
Keeping up with the field's pace isn't about bookmarking — it's about fixed daily/weekly information entry points. The four entries below cover three layers: "papers, news, discussion". Just keep one always open.
Papers with Code: The Bridge Between Papers and Code
One-sentence positioning: The authoritative directory for "does this paper have an open-source implementation, and what ranking does it achieve on which benchmarks".
PwC, originally maintained by Meta AI and later acquired by Hugging Face, connects papers, code, and benchmark leaderboards. Before reading a paper, check PwC first: if there's an implementation, you can reproduce it; if there's a leaderboard, you know what SOTA looks like. Its paper trends page has now been merged into Hugging Face (huggingface.co/papers) for more real-time information.
- Official link: paperswithcode.com
arXiv: The Frontline for Papers
One-sentence positioning: The premiere venue for virtually all AI papers — machine learning papers are almost all under the cs.LG category.
arXiv needs little introduction, but what's worth emphasizing is how to subscribe: cs.LG (machine learning), cs.CL (computational linguistics), and cs.CV (computer vision) — daily digest lists for these three categories are the baseline action for keeping up with the frontier. The subscription method: spend 15 minutes skimming titles each workday, rather than waiting for WeChat public account reposts.
- Official link: arxiv.org | arxiv.org/list/cs.LG/recent
Reddit r/MachineLearning: High-Quality Discussion Hub
One-sentence positioning: The English-language world's most concentrated ML discussion zone; the weekly "three papers" post is the highlight.
r/MachineLearning's unique value lies in its comments — paper authors answer questions directly, practitioners share engineering experience, and the information density far exceeds typical news accounts. Every Wednesday, the "This Week in Machine Learning & AI" post curates three papers worth reading — a low-cost way to stay current.
- Official link: reddit.com/r/MachineLearning
Zhihu: The Main Chinese-Language Discussion Hub
One-sentence positioning: The highest-quality machine learning Q&A community in the Chinese-speaking world — the go-to source for answers to "what is X / how to get started / how to choose a direction" type questions.
Zhihu's machine learning and deep learning topics host many long-form answers accumulated from practitioners and researchers, especially suitable for resolving "conceptual confusion" and "path selection" type questions. Be mindful of timeliness: many "intro advice" posts from before 2020 are outdated; look at both the upvotes and the update date.
- Official link: zhihu.com
The Golden Ratio for Information Intake
Daily: skim arXiv titles (15 minutes); Weekly: read r/MachineLearning's recommended papers + one newsletter (e.g., DeepLearning.AI's The Batch, deeplearning.ai/the-batch); Monthly: pick one highly cited paper from PwC for deep reading. News accounts should be supplementary, not primary sources.
VIII. Chinese-Language Resources
The value of Chinese-language resources lies in "lowering the first barrier", but English proficiency determines how far you can go — this section's positioning is an entry accelerator, not a replacement.
Dive into Deep Learning Chinese Edition (covered above, here is the companion video entry)
One-sentence positioning: China's top deep learning textbook, with Bilibili companion videos + runnable code + discussion community — the complete three-piece set.
Beyond the official website, its GitHub repository (github.com/d2l-ai/d2l-zh) provides all Notebooks and PDFs, and the author has complete Chinese explanation videos on Bilibili. When formulas block you, pair it with the site's Math Primer for on-demand catch-up.
Zhou Zhihua's Machine Learning ("Watermelon Book") & Datawhale's Pumpkin Book
One-sentence positioning: China's most classic machine learning textbook (the Watermelon Book) and its companion open-source formula derivation book (the Pumpkin Book).
Zhou Zhihua's Machine Learning is widely used as a textbook in Chinese universities, characterized by a balance of "machine learning + statistical perspective" and concise explanations — but it skips many formula derivations, making it unfriendly for self-learners. The Pumpkin Book (open-sourced by the Datawhale community) fills in every skipped formula derivation from the Watermelon Book one by one; the two books paired together are the golden combination for learning classic ML in Chinese.
- Official link: github.com/datawhalechina/pumpkin-book
Datawhale: Chinese AI Open-Source Learning Community
One-sentence positioning: China's most active AI open-source learning organization, producing many "Chinese, runnable tutorial repositories".
Beyond the Pumpkin Book, Datawhale also maintains dozens of repositories including Dive into Large Models, Open-Source Large Model Usage Guide (fine-tuning/deployment tutorials), all open-source, organized in Jupyter Notebook format, carrying forward the same style as D2L — an important supply station for Chinese LLM engineering learning.
- Official link: github.com/datawhalechina
Jiqizhixin & Qbitai: The Dual Champions of Chinese AI News
One-sentence positioning: China's two fastest-updating, most comprehensive AI vertical media.
Jiqizhixin (jiqizhixin.com) leans academic and industry-depth, with high-quality paper interpretations; Qbitai (qbitai.com) leans industry dynamics and product news, with fast speed. Both are great entry points for "10 minutes daily to understand what's happening in the industry" — but remember: news is for conversation; courses and papers are for capability — don't get the time allocation backwards.
- Official link: jiqizhixin.com | qbitai.com
How to String Them Together
Finally, a "phase → main resource → site companion page" reference table for on-demand use:
| Your Phase | Main Resource (this page) | Site Companion |
|---|---|---|
| Brand new to ML, don't know what it is | Andrew Ng's first two courses + fast.ai lesson 1 | What Is Machine Learning + Glossary |
| Can run code, want to learn systematically | D2L (deep learning) / ISL (statistical perspective) | Learning Paths: Three Routes |
| Feel math is insufficient | Flower Book chapters 1–5, look up as needed | Math Primer |
| Want to prove you can do it | One Kaggle intro competition + public solution reproduction | Framework Comparison |
| Want to understand a specific direction | CS231n (vision) / CS224n (NLP) / Zero to Hero (under the hood) | Corresponding concept sections |
| Building production systems | MLflow docs + corresponding tool docs | Practice modules |
Remember three iron rules: pick only one main thread; hands-on beats bookmarking; official docs beat secondhand tutorials. Bookmarking 100 links is less effective than completing 1 project — now, close this page, go to Learning Paths, and pick a thread to start.