Theme
Glossary
In one sentence: a quick-reference dictionary of deep learning terms organized by topic — seven groups of 60+ core terms, each with a Chinese name, English name, and one-sentence definition. Commonly confused terms are contrasted in pairs, and deep-dive pages in this handbook are bolded and linked.
Usage: Look up terms you encounter in other sections here for definitions. The distinctions at the end of each group contrast commonly confused pairs. Click bolded links to jump to corresponding pages for deeper understanding. For math notation, see the Math Primer.
1. Fundamentals
| Term | Definition |
|---|---|
| Deep Learning | A machine learning approach using multi-layer neural networks for representation learning. See What is Deep Learning |
| Representation Learning | Letting models automatically learn feature representations from data, rather than manually designing features. See Representation Learning & Pre-training |
| Neural Network | A function approximator composed of stacked layers with learnable parameters. See Neural Network Fundamentals |
| Parameter | Learnable weights and biases within the network. The training objective is to optimize them. |
| Hyperparameter | Parameters set manually before training: learning rate, batch size, number of layers, etc. |
| Gradient | The vector of partial derivatives of the loss w.r.t. parameters, pointing in the direction of steepest loss increase. The opposite direction is the descent direction. See Optimization & Gradient Descent |
| epoch / batch / iteration | epoch = one complete pass through the training set; batch = a group of samples used in one iteration; iteration = one parameter update |
| Loss Function | A function measuring the gap between model predictions and true labels. Training is the minimization of this function. See Loss Functions & Output Layers |
Distinction: epoch vs iteration — epoch is "finishing one complete pass through all data"; iteration is "one parameter update." With 10,000 samples and batch size 500, there are 20 iterations per epoch.
2. Models & Layers
| Term | Definition |
|---|---|
| MLP (Multi-Layer Perceptron) | The most basic feedforward network, composed of stacked fully connected layers |
| CNN (Convolutional Neural Network) | A network using convolutions + pooling to extract local features, with weight sharing. See CNNs & Computer Vision |
| RNN (Recurrent Neural Network) | A network that processes sequences step by step in time, passing hidden states across steps. See RNN & Sequence Modeling |
| LSTM / GRU | RNN variants introducing gating mechanisms to mitigate vanishing gradients |
| Transformer | An architecture based on self-attention that can process sequences in parallel. See Transformer Architecture |
| Attention | A mechanism that aggregates information weighted by relevance, centered on the QKV (query-key-value) mechanism. See Attention Mechanism |
| Embedding | Mapping discrete symbols (words, IDs) to dense vectors as model input representations |
| Positional Encoding | A method to inject sequence position information into the order-independent self-attention mechanism |
| Residual / Skip Connection | A cross-layer direct connection (x + F(x)) that mitigates vanishing gradients in deep networks |
| Normalization Layers (BN/LN) | BatchNorm normalizes across the batch dimension; LayerNorm normalizes across the feature dimension. See Initialization & Normalization |
| Receptive Field | The size of the input region corresponding to a given output neuron |
Distinction: BatchNorm vs LayerNorm — BN depends on batch statistics and is suited for images; LN is independent of batch and suited for sequences / variable-length inputs, which is why Transformer uses LN.
3. Training & Optimization
| Term | Definition |
|---|---|
| Backpropagation | An algorithm that efficiently computes gradients using the chain rule. See Backpropagation & Automatic Differentiation |
| Gradient Descent | An optimization method that iteratively updates parameters in the direction opposite to the gradient |
| SGD (Stochastic Gradient Descent) | Uses a mini-batch of samples to estimate gradients, balancing efficiency with stochasticity |
| Momentum | Accumulates historical gradient direction to accelerate convergence and suppress oscillation |
| Adam / AdamW | Adaptive learning rate optimizers. AdamW corrects Adam's weight decay implementation and has become the standard for pre-training / fine-tuning |
| LR Schedule (Learning Rate Schedule) | Dynamically adjusts learning rate during training: cosine, warmup, step decay, etc. See Training Recipes & Hyperparameter Tuning |
| Gradient Clipping | Scales down gradients when their norm exceeds a threshold, preventing gradient explosion |
| Mixed Precision | FP16/BF16 and FP32 mixed computation, saving VRAM and accelerating training |
| Initialization | Xavier/Kaiming and other methods that set weight initial values per layer scale. See Initialization & Normalization |
| Learning Rate | The step size for each parameter update — the most sensitive hyperparameter |
| Loss Landscape / Local Minima | The geometric structure of high-dimensional loss space. In practice, "saddle points and flat regions" are more common than local minima |
Distinction: GD vs SGD vs mini-batch SGD — GD uses full data to compute one gradient (accurate but slow); SGD strictly updates per sample (noisy); practice uses mini-batch SGD (a trade-off). When people say "SGD," they usually mean mini-batch SGD.
4. Regularization
| Term | Definition |
|---|---|
| Overfitting | The model memorizes training set noise, leading to poor generalization. See Overfitting & Regularization |
| Underfitting | The model has insufficient capacity and can't fit the training set well |
| Dropout | Randomly drops some neurons during training to prevent co-adaptation |
| Weight Decay (L2) | Penalizes large weights, making the model smoother |
| Data Augmentation | Transforms training samples to generate new samples. See Data & Data Engineering |
| Early Stopping | Stop training when the validation metric no longer improves |
| Batch Normalization as Regularization | The noise introduced by BN (batch statistic fluctuations) has a slight regularization effect (debatable among researchers, but this effect shouldn't be relied upon) |
| Generalization | A model's ability to perform on unseen data — the fundamental goal of training |
Distinction: Regularization vs Data Augmentation — Regularization constrains model complexity (modifies the model/loss); data augmentation expands the data distribution (modifies the data). Both can be applied simultaneously without conflict.
5. Generative Models & LLMs
| Term | Definition |
|---|---|
| VAE (Variational Autoencoder) | An encoder-decoder generative model with latent space regularization. See VAE & GAN |
| GAN (Generative Adversarial Network) | Generator and discriminator trained adversarially. See VAE & GAN |
| Diffusion Model | A generative model that learns step-by-step denoising. See Diffusion Models & Generative AI |
| Pre-training | Training general-purpose capabilities on large-scale unlabeled/weakly-labeled data. See Representation Learning & Pre-training |
| Fine-tuning | Continuing training on a pre-trained model using small-scale task-specific data |
| LoRA / PEFT | Low-rank Adaptation / Parameter-Efficient Fine-Tuning: trains only a small number of additional parameters |
| RLHF (Reinforcement Learning from Human Feedback) | Trains a reward model using human preferences, then aligns model behavior via reinforcement learning |
| RAG (Retrieval-Augmented Generation) | Retrieves external knowledge to assist generation, mitigating hallucination and timeliness issues |
| Hallucination | An LLM generates content that seems plausible but is wrong or unsubstantiated |
| Context Window | The maximum number of input tokens a model can process at once |
| Token | The basic unit after text tokenization — Chinese characters typically correspond to ~1–2 tokens |
| KV Cache | Caches the keys and values of already-generated tokens during autoregressive decoding, accelerating inference |
Distinction: Pre-training vs Fine-tuning vs RAG — Pre-training learns general knowledge (high cost); fine-tuning injects tasks/style (moderate cost, modifies parameters); RAG introduces immediate external knowledge (no parameter changes). They solve different problems. See Large Language Models (LLMs).
6. Evaluation & Data
| Term | Definition |
|---|---|
| Accuracy / Precision / Recall / F1 | Core classification metrics. Precision = proportion of predicted positives that are correct; Recall = proportion of actual positives that are found; F1 is the harmonic mean of the two. See Deep Learning Evaluation & Experimentation |
| AUC (Area Under the ROC Curve) | A metric measuring ranking ability, robust to class imbalance |
| Confusion Matrix | A tabular count of predicted vs actual classifications |
| FID (Fréchet Inception Distance) | A metric measuring the distance between generated and real image distributions, used to evaluate GANs/diffusion models |
| Perplexity | A measure of how "surprised" a language model is by text (lower is better) |
| Data Leakage | Training information leaking into the test set, inflating evaluation scores. See Common Pitfalls & Anti-patterns |
| Class Imbalance | When class sample counts are highly skewed — requires resampling or weighted loss. |
| Cross-validation | Splitting data into folds and rotating the validation set — commonly used with small datasets |
| Offline / Online Evaluation | Evaluation on a static dataset vs. evaluation on real production traffic |
Distinction: Precision vs Recall — "What fraction of predicted positives are true?" vs. "What fraction of true positives are found?" Both matter for spam filtering, but recall is more important for medical screening (don't miss diagnoses).
7. Production & Trust
| Term | Definition |
|---|---|
| MLOps | The engineering system for taking models from development to production: experiment tracking, CI/CD, monitoring, rollback. See MLOps & Model Deployment |
| Quantization | Reducing model weights from FP32 to low precision like INT8 to accelerate inference and save VRAM |
| Distillation | Using outputs from a large model (teacher) to guide the training of a smaller model (student) |
| Pruning | Removing unimportant weights/neurons to compress the model |
| Inference | Using a trained model to make predictions, as opposed to training |
| Distributed Training | Cross-device training methods: data parallel, model parallel, tensor parallel, etc. See Training Recipes & Hyperparameter Tuning |
| Interpretability | The ability to understand why a model makes its decisions. See Interpretability & Fairness |
| Fairness | A model's performance shows no systematic bias across different groups |
| Robustness | A model's ability to maintain performance under distribution shift / adversarial perturbations |
| Drift | Model degradation caused by changes in online data distribution over time |
| Benchmark | Standard datasets and metrics used for cross-model capability comparison |
Distinction: Quantization vs Distillation vs Pruning — Quantization changes numerical precision; distillation changes the learning objective (learning from the teacher); pruning changes the structure (removing redundancy). All three can be applied together and form the standard three-step toolkit before model deployment.
Usage Tips
- When you encounter an unfamiliar term on other pages, check this table first for a one-sentence definition, then click the bolded link for deeper understanding;
- Before an interview, go through all seven groups, focusing on memorizing the distinction pairs — interviewers love asking paired comparisons;
- When you discover an unfamiliar term corresponds to a skill keyword in the JD Knowledge Map, immediately add it to your study plan.
Further Reading
- Math Primer — The math definitions behind the terms
- Deep Learning System Anatomy — How components assemble into a complete system
- Learning Paths: Three Routes — Systematically digest terms along your chosen route
- Interview Question Bank — Use terms to ace interview questions
- Common Pitfalls & Anti-patterns — Practical pitfalls behind misused terminology