Skip to content

DEEP LEARNING HANDBOOK

Deep Learning Handbook

A systematic knowledge graph from data to intelligence — neural networks, backpropagation, attention, generative models, large language models, MLOps, and career paths.

If data is ore, deep learning is the entire process that teaches machines to "refine gold" on their own — network design, representation learning, training and optimization, evaluation and iteration, deployment and monitoring. The patterns are no longer hand-engineered; they're learned automatically by multi-layer networks from data. Learn more →

Where to start

50+ pages aren't a library to read cover to cover — they're a roadmap you can assemble on demand

✨ Editor's Picks

Hand-picked from each section — these are the ones you should read first

Browse All

Browse by section, or search directly

✨ Recommended

🧠 Backpropagation and Automatic Differentiation

Backpropagation is the foundation of deep learning training: the chain rule + gradient backpropagation. This article starts from "why we need gradients," manually derives a small network step by step, then transitions to the two modes of automatic differentiation, and finally covers the correct usage of PyTorch autograd, common pitfalls, and the root cause of vanishing/exploding gradients.

backpropagationautomatic differentiationchain rulePyTorchgradients
✨ Recommended

🧠 Overfitting and Regularization

The nemesis of deep learning models is not "insufficient capability" but "memorizing too much." This article uses the bias-variance decomposition to clarify overfitting mechanisms, systematically covers L1/L2, Dropout, early stopping, data augmentation, label smoothing, EMA, Mixup, and other regularization weapons, and provides engineering experience on combination strategies and intensity selection.

overfittingregularizationDropoutdata augmentationgeneralization

🧠 Representation Learning and Pretraining

The secret weapon of deep learning isn't just fitting ability — it's "learning good representations." This article clarifies feature hierarchies and end-to-end thinking, the semantic structure of embedding spaces, and surveys transfer learning and fine-tuning strategies, the pretrain-finetune paradigm, and goes deep into self-supervised learning (SimCLR/MAE) and representation evaluation methods.

representation learningpretrainingtransfer learningself-supervised learningembeddings
✨ Recommended

🔍 CNN and Computer Vision

Convolutional Neural Networks were the first breakthrough that ignited deep learning. This article breaks down the three core building blocks—convolution, pooling, and residual connections—along with the architectural evolution from LeNet to ViT, covers the family of vision tasks, the transfer learning recipe (with PyTorch code) and training tricks, and surveys the current state of large vision models like CLIP, SAM, and VLMs.

CNNComputer VisionTransfer LearningImage ClassificationViT

🔍 Deep Learning Recommender Systems

Recommender systems are the most profitable battlefield for deep learning in internet industry. This article covers the four-layer architecture of recall/ranking/re-ranking, from matrix factorization to neural collaborative filtering, two-tower recall, SASRec sequential recommendation, DCN/Wide&Deep CTR ranking, PinSage graph-based recommendation, and MMoE multi-objective learning, with practical perspectives on cold start, evaluation, and deployment.

recommender systemsctr predictiontwo-tower modelsrecallmulti-objective learning
✨ Recommended

🔍 Diffusion Models and Generative AI

Diffusion models sidestep the blurriness of VAEs and the instability of GANs through a "forward noise-adding, backward noise-removing" process, becoming the dominant engine for text-to-image, video, audio, and other generative AI tasks. This article unpacks DDPM principles, DDIM accelerated sampling, Classifier-Free Guidance, the Stable Diffusion stack, and LoRA/ControlNet, and discusses evaluation and limitations.

diffusion modelsgenerative AIstable diffusionddpmlora

🔍 Graph Neural Networks

Graph neural networks extend deep learning to structured data on nodes and edges. This article unpacks graph representations, the message passing framework, and three mainstream models — GCN, GAT, and GraphSAGE — covering node classification, link prediction, and graph classification tasks and applications (recommenders, molecules, physics simulation), while confronting limitations like neighbor sampling, over-smoothing, and scalability.

graph neural networksgcngatmessage passinggraph representation learning
✨ Recommended

🔍 Transformer Architecture

The Transformer replaced recurrence with self-attention, becoming the de facto standard in the era of large models. This article breaks down the Encoder/Decoder architecture, multi-head self-attention, positional encoding, Pre-Norm/Post-Norm, provides training tricks like warmup, surveys the BERT vs GPT families and ViT/cross-modal extensions, and covers engineering optimizations like FlashAttention and KV cache.

TransformerSelf-AttentionBERTGPTPositional Encoding

🔍 VAE and GAN

VAEs and GANs are the two foundational paradigms of deep generative modeling: VAEs learn latent spaces through variational inference, GANs generate realistic data through adversarial games. This article breaks down the reparameterization trick and ELBO, minimax games and mode collapse, surveys the improvement lineage from DCGAN/WGAN-GP/StyleGAN and conditional generation, and compares the current state against diffusion models.

VAEGANGenerative ModelsLatent SpaceAdversarial Training

📄 Frontier Advances

Important breakthroughs and trends in deep learning from recent years: scaling laws and synthetic data, MoE sparse experts, multimodal unification, RLHF and alignment (DPO/RLVR), diffusion application extension, test-time compute (o1-class), Agents and tool use, efficiency revolution, and scientific discovery. Each section covers what it is, why it matters, representative work, and linked pages, with guidance on how to read time-sensitive data.

frontier advanceslarge modelsdiffusion modelsalignmentefficiency revolution
✨ Recommended

🛠️ Training Recipes and Hyperparameter Tuning

Deep learning training is a blend of following proven recipes and exercising judgment in the moment. This article provides default recipe tables for data, initialization, loss, optimizer, learning rate, regularization, and normalization, explains learning rate range testing, batch size interplay, warmup and cosine annealing, transfer learning and memory management, and ends with experiment reproducibility discipline.

practicehyperparametersoptimizerlearning-ratereproducibility