Skip to content

Glossary

Quick overview Core machine learning glossary: accurate definitions and distinctions for 60+ terms across foundational concepts, data, models & algorithms, deep learning, generative models & LLMs, evaluation & metrics, and production & trusted ML.

Glossary ​

This glossary collects core terms from the machine learning field, organized into seven groups by topic: foundational concepts, data, models & algorithms, deep learning, generative models & LLMs, evaluation & metrics, and production & trusted ML. Each term provides a Chinese name, English name, and definition; easily confused term pairs have separate distinction notes.

How to use

Terms don't need to be read in order. Start by reading "Foundational Concepts" to build a vocabulary framework, then return to look up unfamiliar words when you encounter them in other sections. Bolded internal links point to in-depth discussion pages for that term.

Foundational Concepts ​

Machine Learning (ML) Technology that enables computers to automatically discover patterns from data and use them for prediction/decision-making. The core is "generating rules from data" rather than "humans writing rules." Four elements: task T, experience E, performance P, and improvement mechanism. See What Is Machine Learning.

Artificial Intelligence (AI) The overarching goal of enabling machines to perform tasks requiring human intelligence. Machine learning is a subfield of AI (currently the most successful one); rule-based systems, knowledge graphs, etc. also belong to AI but are not ML. See ML vs AI.

Deep Learning (DL) A subfield of machine learning that uses multi-layer neural networks for representation learning. The core is end-to-end feature learning: shallow layers learn low-level features, deep layers learn high-level semantics. See Deep Learning Fundamentals.

Data Science (DS) A complete-process discipline for answering business questions with data and supporting decisions: collection and cleaning → analysis → modeling → communication. Machine learning is the modeling step within it.

Supervised Learning A paradigm that uses labeled data $(x, y)$ to learn an $x \to y$ mapping, covering classification, regression, and ranking. See Supervised Learning.

Unsupervised Learning A paradigm that uses only unlabeled data $x$ to discover data structures: clustering, dimensionality reduction, anomaly detection, association rules. See Unsupervised Learning.

Reinforcement Learning (RL) A paradigm where an agent learns an optimal decision sequence through interaction with an environment, guided by a reward signal. Key challenges are credit assignment and the exploration-exploitation tradeoff. See Reinforcement Learning.

Feature A measurable attribute of the model input, i.e., "the numerical representation that raw data is transformed into." Feature quality determines the model's upper bound.

Label The ground truth $y$ in supervised learning, paired with input $x$ to form training samples.

Generalization A model's ability to perform well on unseen data. Everything ML aims at can be reduced to "generalization" — performing well on the training set does not equal learning.

Overfitting The model learns noise in the training data as if it were a pattern, resulting in good training performance but poor generalization. See Overfitting and Regularization.

Underfitting The model has insufficient capacity and performs poorly even on the training data (high bias). Opposite of overfitting.

Bias-Variance Tradeoff Prediction error = bias² + variance + noise. The more complex a model, the lower the bias and higher the variance; optimal complexity lies in the middle. See Model Evaluation and Validation.

Regularization A family of techniques that "prevent the model from learning noise": L1/L2, Dropout, early stopping, data augmentation. The essence is "adding a tax on model complexity".

Model The parameterized function (e.g., $w$, $b$ for linear regression) obtained by training a learning algorithm on data — it is the vehicle for "the rules that generate data".

Algorithm The mathematical procedure of the learning process (e.g., gradient descent, EM) responsible for solving model parameters from data.

Hyperparameter Parameters set before training (learning rate, tree depth, $\lambda$), not learned from data but determined by searching on the validation set. See Hyperparameter Tuning Practice.

Data ​

Training / Validation / Test Set Three data splits: training set for fitting parameters, validation set for selecting hyperparameters, test set for final evaluation. The test set must never be touched during any iteration.

Data Leakage Test set information leaking into the training process, leading to artificially high offline metrics and catastrophic online performance. Six sources: global statistics, future information, target leakage, duplicated samples, group leakage, augmentation leakage. See Data and Data Engineering.

Class Imbalance Disproportionate positive-to-negative sample ratios, causing accuracy to be misleading and the model to bias toward the majority class. Countermeasures: resampling, class weights, changing metrics.

Data Augmentation Using prior knowledge to create more training samples (image flipping/cropping, noise injection) — a "free data expansion" regularization technique.

Feature Engineering The full process of transforming raw data into numerical features that models can effectively utilize: missing values, encoding, scaling, combinations, and selection. See Feature Engineering.

Feature Store A system that unifies the definition and computation of training/inference features, solving the engineering disaster of "training features and online features being inconsistent". See MLOps.

Data Drift Model performance degradation caused by the online input feature distribution changing over time. Detection: PSI, KS test.

Concept Drift The $x \to y$ mapping relationship itself changes (e.g., policy changes alter risk control patterns), making it harder to detect than data drift.

Models & Algorithms ​

Linear Regression A regression model assuming $y = w \cdot x + b$, solved via least squares / gradient descent. The default baseline for all regression tasks. See Linear Models.

Logistic Regression Linear output passed through sigmoid to get probability + cross-entropy loss for classification. The name contains "regression" but it's actually the most important classification baseline.

Decision Tree An if-then rule tree that recursively partitions samples by feature values. Interpretable, normalization-free, and natively handles categorical features.

Random Forest Bagging ensemble of multiple decision trees: each tree is trained on a random subsample + random feature subset, outputting via voting/averaging. Robust, resistant to overfitting.

Gradient Boosting Additive ensemble that fits residuals tree by tree: GBDT → XGBoost → LightGBM. The de facto king of tabular data. See Tree Models and Ensemble Learning.

Support Vector Machine (SVM) Uses kernel functions to map data to high-dimensional space and find the maximum-margin hyperplane. Still valuable for small-sample, high-dimensional scenarios; has been replaced by deep learning for unstructured data.

K-Nearest Neighbors (KNN) Finds the K closest samples by distance for voting/averaging. Simple but slow at prediction (requires distance computation), fails in high dimensions.

Clustering Unsupervised grouping of similar samples: K-Means, DBSCAN, hierarchical clustering, GMM. See Unsupervised Learning.

Dimensionality Reduction Mapping high-dimensional data to lower dimensions while preserving major structure: PCA (linear), t-SNE/UMAP (non-linear visualization), autoencoders.

Principal Component Analysis (PCA) Projecting data onto orthogonal directions of maximum variance. The default choice for linear dimensionality reduction; requires standardization before use.

Anomaly Detection Finding points significantly different from the majority: Isolation Forest, LOF, One-Class SVM, autoencoder reconstruction error.

Ensemble Learning Combining multiple models to improve performance: bagging (parallel, reduces variance), boosting (sequential, reduces bias), stacking (meta-learning).

Bayesian Methods A statistical framework treating parameters as random variables, using prior + likelihood to compute posterior: Naive Bayes, Gaussian Processes, Bayesian Optimization.

Deep Learning ​

Neural Network A network of neurons (weighted sum + activation function) connected in layers, trained via backpropagation. See Deep Learning Fundamentals.

Backpropagation The algorithm that uses the chain rule to propagate the loss's gradient to each layer's parameters, combined with gradient descent for parameter updates. The engine of deep learning training.

Activation Function Functions that introduce nonlinearity (ReLU/GELU/sigmoid/softmax), the mathematical prerequisite for "why deep networks are useful".

Convolutional Neural Network (CNN) Networks designed for images: convolutional kernels slide to extract local features, parameter sharing, and pooling for compression. Residual connections allow them to reach 150+ layers. See CNN.

Recurrent Neural Network (RNN) Networks that process sequences step by step in time; LSTM uses gating to solve long-term dependencies. Replaced by Transformers in the 2020s.

Attention A mechanism for aggregating information by "relevance weights" when processing sequences: Q/K/V three vectors, weights = softmax(Q·Kᵀ/√d). See Transformer.

Transformer A sequence architecture based entirely on attention, abandoning recurrence (2017). Strong parallelism, excellent long-range dependency modeling — the foundation of BERT/GPT/all LLMs.

Self-Attention Attention within a sequence where every token interacts pairwise: every word attends to other words in its context. Multi-head attention captures different relationships in parallel.

Positional Encoding A mechanism that supplements sequential information to the parallel-computing Transformer (sinusoidal, learnable, RoPE).

Residual Connection The skip-connection structure output = F(x) + x, solving deep network degradation/training difficulty issues — a standard component in modern deep networks.

Normalization Layer Techniques that stabilize training: BatchNorm (along batch dimension, used in CNNs), LayerNorm (along feature dimension, used in Transformers).

Embedding Mapping discrete objects (words, users, items) into dense vector representations, where semantically similar objects have similar vectors.

Transfer Learning Transferring knowledge from pre-trained models (ImageNet/BERT) to a target task: feature extraction or fine-tuning. The core strategy for small-data projects.

Fine-Tuning Continuing to train on task data with a small learning rate on top of a pre-trained model.

Loss Function A function measuring the gap between prediction and ground truth (MSE/cross-entropy/Huber), the optimization target of training. Matching the task's output distribution is key.

Optimizer Concrete implementation of gradient descent: SGD, Momentum, Adam, AdamW. The default for deep learning is AdamW, standard for large models. See Optimization.

Learning Rate The step size at each step of gradient descent, the most important hyperparameter. Too large: oscillation and divergence. Too small: turtle-speed convergence.

Vanishing/Exploding Gradient In deep networks, gradients either vanish layer by layer (shallow layers don't update) or explode (parameters diverge). Countermeasures: ReLU, residuals, normalization, gradient clipping.

Dropout A regularization technique that randomly drops some neurons during training, forcing the network to learn redundant representations — equivalent to "ensemble within a single model".

Early Stopping Monitor validation loss and stop training when it stops decreasing, keeping the best model. Free regularization, enabled by default in deep learning.

Quantization Reducing model weights from FP32 to FP16/INT8/INT4, halving memory usage and accelerating inference with minimal precision loss. The top choice for deployment optimization.

Distillation Using a large model (teacher) to teach a small model (student), enabling the small model to approximate the large model's performance. A common technique for cost reduction and efficiency improvement.

Generative Models & LLMs ​

Generative Model A model that learns data distribution $P(x)$ and can sample new samples from it: VAE, GAN, diffusion models, autoregressive models. See Generative Models.

GAN (Generative Adversarial Network) Adversarial training between a generator and discriminator: G creates fakes to deceive D, D distinguishes real from fake. Produces sharp images but training is unstable, prone to mode collapse.

Diffusion Model First adds noise to data until it becomes pure noise, then learns to progressively denoise and reconstruct. The current de facto standard for text-to-image/audio-to-image (Stable Diffusion, DALL·E). See Diffusion Models.

Autoregressive Generation A generation method that predicts the next token step by step — the core of the GPT series. See Large Language Models.

Large Language Model (LLM) A generative model based on Transformer decoders, pre-trained on ultra-large-scale corpora, with emergent abilities in understanding, reasoning, and conversation.

Pretraining The phase of training foundational capabilities (language/image representation) on massive unlabeled corpora. Produces foundation models.

Supervised Fine-Tuning (SFT) Fine-tuning with "instruction-response" pairs to teach the model to follow instructions.

RLHF (Reinforcement Learning from Human Feedback) A pipeline that trains a reward model using human preferences, then aligns the model's behavior using reinforcement learning (PPO). The key technology behind ChatGPT.

DPO (Direct Preference Optimization) A simplified alignment method that directly fine-tunes based on preferences, without training a reward model — the mainstream approach in the open-source community.

In-Context Learning The ability to complete tasks without updating weights, relying solely on input prompts (zero-shot/few-shot examples).

Chain-of-Thought (CoT) Prompting the model to "think step by step before answering", significantly boosting performance on reasoning tasks.

Retrieval-Augmented Generation (RAG) A paradigm of "retrieve first, then generate": vector retrieval injects private/real-time knowledge into context, addressing three major shortcomings: outdated knowledge, hallucinations, and lack of private data.

LoRA (Low-Rank Adaptation) A fine-tuning method that trains only low-rank adaptation matrices while freezing original weights — can run on a single GPU, the de facto standard for fine-tuning.

Hallucination Generative models fabricating non-existent "facts" — they learn "sequences that sound human" rather than "fact databases". Mitigation: RAG, citations, human review.

Prompt Engineering The technique of designing input prompts (instructions, examples, format constraints) to guide model output quality.

Temperature A sampling parameter that controls generation randomness: low temperature is more deterministic (factual tasks), high temperature is more diverse (creative tasks).

Token The basic unit of text processing in a model (~0.7 English words / ~1 Chinese character); billing and context windows are measured in tokens.

Evaluation & Metrics ​

Accuracy The proportion of correct predictions. A trap metric under class imbalance (guessing the majority class entirely still yields high accuracy).

Precision / Recall Precision = the true-positive proportion among predicted positives (penalizes false positives); Recall = the proportion of actual positives that are recovered (penalizes false negatives). Which one to prioritize depends on "which type of error is more costly".

F1 The harmonic mean of precision and recall, a compromise metric when both matter equally.

Confusion Matrix A four-cell table of TP/FP/FN/TN, the starting point for all classification metrics.

ROC / AUC ROC curve area = the probability that a random positive scores higher than a random negative. Threshold-independent, relatively insensitive to imbalance.

Mean Squared Error (MSE) A regression metric, squares penalties, sensitive to outliers, well-behaved differentiability.

R² (R-squared / Coefficient of Determination) A regression metric measuring "what fraction of variance the model explains", compared against a mean baseline; can be negative.

Cross-Validation K-fold alternating training/validation, reporting mean and variance — more reliable than a single holdout on small data. Variants: StratifiedKFold, GroupKFold, TimeSeriesSplit.

A/B Testing Randomly splitting users into two groups online, comparing the real business impact of old vs. new models/versions. The final arbiter of offline metrics.

pass@k A probability metric measuring "the probability of at least one success in k samples", assessing the model's capability ceiling (commonly used in LLM evaluation).

Production & Trusted ML ​

MLOps (Machine Learning Operations) Engineering practices for transforming ML systems from notebooks into production-stable services (experiment tracking, model registry, deployment, monitoring, retraining). See MLOps.

Experiment Tracking Recording the code/data/hyperparameters/metrics of each experiment to ensure reproducibility. Tools: MLflow, W&B.

Model Registry Model versioning, state management, approval for deployment, and rollback capability. Tool: MLflow Registry.

Online / Batch Inference Real-time API response vs. scheduled batch scoring. Choose online for latency-sensitive tasks, batch for non-time-critical (saves ~90% cost).

Interpretability Understanding "why the model made this decision": global (feature importance), local (SHAP, LIME), transparent models (linear/tree). See Interpretability and Fairness.

SHAP A method for explaining model outputs based on game-theoretic Shapley values — the local explanation standard. Computationally expensive, implemented via approximations.

Fairness Whether a model treats different groups equally: demographic parity, equal opportunity, equalized calibration (these three cannot be satisfied simultaneously; requires business definition). Bias primarily comes from data.

Prompt Injection An attacker hides malicious instructions in content the model will read (webpages, documents, tool outputs), inducing the model to deviate from its intended task. A security threat specific to LLM applications.

Guardrail Programmatic checks and interceptions on model inputs and outputs: content filtering, format validation, action allowlists.

Confusable Terms Distinguished ​

Machine Learning / Deep Learning / Reinforcement Learning

Machine learning is the umbrella term for "learning patterns from data"; deep learning is the branch within it using multi-layer neural networks (images/text/speech); reinforcement learning is a third paradigm (learning decision sequences), an orthogonal dimension to "deep learning" — deep RL (DQN/PPO) is what combines both. See ML vs AI.

Accuracy / Precision / Recall

Accuracy looks at the overall (proportion guessed correctly); precision looks at the credibility of predicted positives (penalizes false positives); recall looks at the recovery rate of actual positives (penalizes false negatives). Business example: spam detection prioritizes precision (don't misflag important emails), cancer screening prioritizes recall (don't miss patients). Metric selection in Model Evaluation and Validation.

Bias / Variance / Overfitting / Underfitting

Underfitting = high bias (model is too simple to learn); overfitting = high variance (model is too complex and memorized the noise). These two are opposite ends of the same tradeoff axis; regularization is the knob that adjusts the pointer back from the "overfitting end".

Fine-tuning / RAG / Prompt Engineering

These three are the toolkit for LLM deployment, each solving different problems: prompt engineering changes "input" (zero cost, simple tasks); RAG adds "knowledge" (factual/private/real-time problems); fine-tuning changes "behavior" (style/format/domain behavior). Use RAG for knowledge issues, fine-tuning for behavior issues — using fine-tuning to feed knowledge is a common misconception.

Data Drift / Concept Drift / Out-of-Distribution

Data drift: input distribution changed (e.g., age distribution shifted); Concept drift: the underlying relationship changed (e.g., feature → outcome relationship broke down); Out-of-distribution (OOD): encountering input types never seen during training. All three are causes of "model obsolescence"; monitoring and retraining are the countermeasures.

Further Reading ​