Skip to content

Glossary

Quick overview A quick-reference dictionary of core deep learning terms: seven groups of 60+ terms covering fundamentals, models & layers, training & optimization, regularization, generative models & LLMs, evaluation & data, and production & trust. Includes distinctions of commonly confused terms and deep-dive links throughout the handbook.

Glossary ​

In one sentence: a quick-reference dictionary of deep learning terms organized by topic — seven groups of 60+ core terms, each with a Chinese name, English name, and one-sentence definition. Commonly confused terms are contrasted in pairs, and deep-dive pages in this handbook are bolded and linked.

Usage: Look up terms you encounter in other sections here for definitions. The distinctions at the end of each group contrast commonly confused pairs. Click bolded links to jump to corresponding pages for deeper understanding. For math notation, see the Math Primer.

1. Fundamentals ​

TermDefinition
Deep LearningA machine learning approach using multi-layer neural networks for representation learning. See What is Deep Learning
Representation LearningLetting models automatically learn feature representations from data, rather than manually designing features. See Representation Learning & Pre-training
Neural NetworkA function approximator composed of stacked layers with learnable parameters. See Neural Network Fundamentals
ParameterLearnable weights and biases within the network. The training objective is to optimize them.
HyperparameterParameters set manually before training: learning rate, batch size, number of layers, etc.
GradientThe vector of partial derivatives of the loss w.r.t. parameters, pointing in the direction of steepest loss increase. The opposite direction is the descent direction. See Optimization & Gradient Descent
epoch / batch / iterationepoch = one complete pass through the training set; batch = a group of samples used in one iteration; iteration = one parameter update
Loss FunctionA function measuring the gap between model predictions and true labels. Training is the minimization of this function. See Loss Functions & Output Layers

Distinction: epoch vs iteration — epoch is "finishing one complete pass through all data"; iteration is "one parameter update." With 10,000 samples and batch size 500, there are 20 iterations per epoch.

2. Models & Layers ​

TermDefinition
MLP (Multi-Layer Perceptron)The most basic feedforward network, composed of stacked fully connected layers
CNN (Convolutional Neural Network)A network using convolutions + pooling to extract local features, with weight sharing. See CNNs & Computer Vision
RNN (Recurrent Neural Network)A network that processes sequences step by step in time, passing hidden states across steps. See RNN & Sequence Modeling
LSTM / GRURNN variants introducing gating mechanisms to mitigate vanishing gradients
TransformerAn architecture based on self-attention that can process sequences in parallel. See Transformer Architecture
AttentionA mechanism that aggregates information weighted by relevance, centered on the QKV (query-key-value) mechanism. See Attention Mechanism
EmbeddingMapping discrete symbols (words, IDs) to dense vectors as model input representations
Positional EncodingA method to inject sequence position information into the order-independent self-attention mechanism
Residual / Skip ConnectionA cross-layer direct connection (x + F(x)) that mitigates vanishing gradients in deep networks
Normalization Layers (BN/LN)BatchNorm normalizes across the batch dimension; LayerNorm normalizes across the feature dimension. See Initialization & Normalization
Receptive FieldThe size of the input region corresponding to a given output neuron

Distinction: BatchNorm vs LayerNorm — BN depends on batch statistics and is suited for images; LN is independent of batch and suited for sequences / variable-length inputs, which is why Transformer uses LN.

3. Training & Optimization ​

TermDefinition
BackpropagationAn algorithm that efficiently computes gradients using the chain rule. See Backpropagation & Automatic Differentiation
Gradient DescentAn optimization method that iteratively updates parameters in the direction opposite to the gradient
SGD (Stochastic Gradient Descent)Uses a mini-batch of samples to estimate gradients, balancing efficiency with stochasticity
MomentumAccumulates historical gradient direction to accelerate convergence and suppress oscillation
Adam / AdamWAdaptive learning rate optimizers. AdamW corrects Adam's weight decay implementation and has become the standard for pre-training / fine-tuning
LR Schedule (Learning Rate Schedule)Dynamically adjusts learning rate during training: cosine, warmup, step decay, etc. See Training Recipes & Hyperparameter Tuning
Gradient ClippingScales down gradients when their norm exceeds a threshold, preventing gradient explosion
Mixed PrecisionFP16/BF16 and FP32 mixed computation, saving VRAM and accelerating training
InitializationXavier/Kaiming and other methods that set weight initial values per layer scale. See Initialization & Normalization
Learning RateThe step size for each parameter update — the most sensitive hyperparameter
Loss Landscape / Local MinimaThe geometric structure of high-dimensional loss space. In practice, "saddle points and flat regions" are more common than local minima

Distinction: GD vs SGD vs mini-batch SGD — GD uses full data to compute one gradient (accurate but slow); SGD strictly updates per sample (noisy); practice uses mini-batch SGD (a trade-off). When people say "SGD," they usually mean mini-batch SGD.

4. Regularization ​

TermDefinition
OverfittingThe model memorizes training set noise, leading to poor generalization. See Overfitting & Regularization
UnderfittingThe model has insufficient capacity and can't fit the training set well
DropoutRandomly drops some neurons during training to prevent co-adaptation
Weight Decay (L2)Penalizes large weights, making the model smoother
Data AugmentationTransforms training samples to generate new samples. See Data & Data Engineering
Early StoppingStop training when the validation metric no longer improves
Batch Normalization as RegularizationThe noise introduced by BN (batch statistic fluctuations) has a slight regularization effect (debatable among researchers, but this effect shouldn't be relied upon)
GeneralizationA model's ability to perform on unseen data — the fundamental goal of training

Distinction: Regularization vs Data Augmentation — Regularization constrains model complexity (modifies the model/loss); data augmentation expands the data distribution (modifies the data). Both can be applied simultaneously without conflict.

5. Generative Models & LLMs ​

TermDefinition
VAE (Variational Autoencoder)An encoder-decoder generative model with latent space regularization. See VAE & GAN
GAN (Generative Adversarial Network)Generator and discriminator trained adversarially. See VAE & GAN
Diffusion ModelA generative model that learns step-by-step denoising. See Diffusion Models & Generative AI
Pre-trainingTraining general-purpose capabilities on large-scale unlabeled/weakly-labeled data. See Representation Learning & Pre-training
Fine-tuningContinuing training on a pre-trained model using small-scale task-specific data
LoRA / PEFTLow-rank Adaptation / Parameter-Efficient Fine-Tuning: trains only a small number of additional parameters
RLHF (Reinforcement Learning from Human Feedback)Trains a reward model using human preferences, then aligns model behavior via reinforcement learning
RAG (Retrieval-Augmented Generation)Retrieves external knowledge to assist generation, mitigating hallucination and timeliness issues
HallucinationAn LLM generates content that seems plausible but is wrong or unsubstantiated
Context WindowThe maximum number of input tokens a model can process at once
TokenThe basic unit after text tokenization — Chinese characters typically correspond to ~1–2 tokens
KV CacheCaches the keys and values of already-generated tokens during autoregressive decoding, accelerating inference

Distinction: Pre-training vs Fine-tuning vs RAG — Pre-training learns general knowledge (high cost); fine-tuning injects tasks/style (moderate cost, modifies parameters); RAG introduces immediate external knowledge (no parameter changes). They solve different problems. See Large Language Models (LLMs).

6. Evaluation & Data ​

TermDefinition
Accuracy / Precision / Recall / F1Core classification metrics. Precision = proportion of predicted positives that are correct; Recall = proportion of actual positives that are found; F1 is the harmonic mean of the two. See Deep Learning Evaluation & Experimentation
AUC (Area Under the ROC Curve)A metric measuring ranking ability, robust to class imbalance
Confusion MatrixA tabular count of predicted vs actual classifications
FID (Fréchet Inception Distance)A metric measuring the distance between generated and real image distributions, used to evaluate GANs/diffusion models
PerplexityA measure of how "surprised" a language model is by text (lower is better)
Data LeakageTraining information leaking into the test set, inflating evaluation scores. See Common Pitfalls & Anti-patterns
Class ImbalanceWhen class sample counts are highly skewed — requires resampling or weighted loss.
Cross-validationSplitting data into folds and rotating the validation set — commonly used with small datasets
Offline / Online EvaluationEvaluation on a static dataset vs. evaluation on real production traffic

Distinction: Precision vs Recall — "What fraction of predicted positives are true?" vs. "What fraction of true positives are found?" Both matter for spam filtering, but recall is more important for medical screening (don't miss diagnoses).

7. Production & Trust ​

TermDefinition
MLOpsThe engineering system for taking models from development to production: experiment tracking, CI/CD, monitoring, rollback. See MLOps & Model Deployment
QuantizationReducing model weights from FP32 to low precision like INT8 to accelerate inference and save VRAM
DistillationUsing outputs from a large model (teacher) to guide the training of a smaller model (student)
PruningRemoving unimportant weights/neurons to compress the model
InferenceUsing a trained model to make predictions, as opposed to training
Distributed TrainingCross-device training methods: data parallel, model parallel, tensor parallel, etc. See Training Recipes & Hyperparameter Tuning
InterpretabilityThe ability to understand why a model makes its decisions. See Interpretability & Fairness
FairnessA model's performance shows no systematic bias across different groups
RobustnessA model's ability to maintain performance under distribution shift / adversarial perturbations
DriftModel degradation caused by changes in online data distribution over time
BenchmarkStandard datasets and metrics used for cross-model capability comparison

Distinction: Quantization vs Distillation vs Pruning — Quantization changes numerical precision; distillation changes the learning objective (learning from the teacher); pruning changes the structure (removing redundancy). All three can be applied together and form the standard three-step toolkit before model deployment.

Usage Tips ​

  1. When you encounter an unfamiliar term on other pages, check this table first for a one-sentence definition, then click the bolded link for deeper understanding;
  2. Before an interview, go through all seven groups, focusing on memorizing the distinction pairs — interviewers love asking paired comparisons;
  3. When you discover an unfamiliar term corresponds to a skill keyword in the JD Knowledge Map, immediately add it to your study plan.

Further Reading ​

References ​