Theme
Unsupervised Learning
Concept Definition: Learning without the Answer Key
Supervised learning learns from "labeled (x, y) pairs," while unsupervised learning deals with data where there's only input x, no label y. The goal isn't to predict an answer, but to discover the inherent structure of the data: which samples are similar? How many natural groups does the data have? Which dimensions matter? Which points are anomalous?
A one-line summary of the four major tasks:
- Clustering: automatically group similar samples;
- Dimensionality Reduction: compress high-dimensional data into lower dimensions while preserving the main structure;
- Anomaly Detection: find points that differ significantly from the majority of samples;
- Association: discover co-occurrence patterns between features/items.
The value of unsupervised learning is often underestimated in real business: in reality, labels are scarce and expensive (annotation costs money and people), while unsupervised learning can directly extract value from massive unlabeled data — customer segmentation, data compression, fraud detection, and feature learning. In the deep learning era, unsupervised "representation learning" (pre-training) has even become the core of large models — see Generative Models and Large Language Models.
Clustering: Grouping Data
Representative Algorithms
| Algorithm | Core Idea | Characteristics |
|---|---|---|
| K-Means | Iterate: assign samples to nearest centroid → update centroid | Fast, simple, requires specifying K, sensitive to outliers/non-spherical distributions |
| DBSCAN | Density-connected samples form clusters | No need to specify K, finds arbitrary shapes, auto-detects noise points |
| Agglomerative Clustering | Bottom-up merging of nearest clusters, produces a dendrogram | Visualization-friendly, view at any granularity |
| GMM (Gaussian Mixture) | Assumes data is a mixture of K Gaussian distributions | Soft clustering (gives probabilities), can model elliptical clusters |
The Full K-Means Pipeline
text
1. Select K initial centroids
2. Loop until convergence:
a. Assign each sample to the nearest centroid (distance is typically Euclidean)
b. Centroid of each cluster = mean of samples in that clusterpython
from sklearn.cluster import KMeans
import numpy as np
X = np.random.default_rng(0).normal(size=(500, 2)) # 500 two-dimensional samples
kmeans = KMeans(n_clusters=3, n_init=10, random_state=42)
labels = kmeans.fit_predict(X)
print(f"Centroids: {kmeans.cluster_centers_}")How to Choose K?
- Elbow method: plot the "intra-cluster sum of squares (inertia) vs. K" curve and look for the bend point;
- Silhouette Score: measures "within-cluster compactness + between-cluster separation," pick the K with the highest score;
- Business-driven: whether K=5 or K=8 for customer segmentation ultimately depends on whether the resulting groups are meaningful and interpretable in the business context.
The Evaluation Challenge of Clustering
Clustering has no "answer key," and evaluation is inherently difficult: internal metrics (silhouette score, inertia) only look at structure and don't guarantee business relevance; external metrics (ARI, NMI) require true labels — but when true labels exist, you'd typically just use supervised learning. Practical advice: prioritize business interpretability, use internal metrics only as a reference.
Dimensionality Reduction: Compressing Data while Preserving Structure
High-dimensional data has three problems: expensive computation, impossible visualization, and the curse of dimensionality (in high-dimensional space, samples become increasingly "sparse," making distance concepts meaningless). Dimensionality reduction maps data into a lower-dimensional space, divided into two categories:
Linear Dimensionality Reduction
PCA (Principal Component Analysis): finds orthogonal directions of maximum variance (principal components) and projects data onto them. Principle: the first principal component direction = the direction of maximum data variance. Uses of PCA:
- Visualization (reduce to 2D/3D);
- Denoising (discard components with low variance);
- Feature compression (prevent overfitting);
- Whitening/preprocessing (works with scaling).
python
from sklearn.decomposition import PCA
pca = PCA(n_components=2) # reduce to 2 dimensions
X_2d = pca.fit_transform(X)
print(f"Explained variance ratio of two PCs: {pca.explained_variance_ratio_}")Choosing the number of components: look at the "cumulative explained variance ratio" curve and pick the number that reaches 80%~95%.
Nonlinear Dimensionality Reduction
| Algorithm | Core Idea | Characteristics |
|---|---|---|
| t-SNE | Preserves high-dimensional neighborhood relations, uses t-distribution for low-dim projection | Great for visualization, but doesn't preserve global structure, results are stochastic and irreversible |
| UMAP | Manifold learning + topological ideas | Faster than t-SNE, preserves global structure better, the go-to choice for visualization in recent years |
| Autoencoder | Neural network: compress to a bottleneck then reconstruct | Nonlinear representation learning, differentiable, can be fine-tuned |
Three Misuses of Dimensionality Reduction
- Must standardize before PCA (otherwise features with large scales dominate the principal components);
- t-SNE distances are not trustworthy — it only preserves relative neighborhood relationships; "far" on the plot doesn't mean truly far;
- Dimensionality reduction for visualization ≠ for modeling: reducing X to 2 dimensions and then training a classifier usually loses information; use PCA/autoencoder, not t-SNE, for dimensionality reduction before modeling.
Anomaly Detection: Finding the "Odd Ones Out"
Business scenarios: fraud detection, industrial quality inspection, network intrusion, data cleaning.
| Method | Idea | Characteristics |
|---|---|---|
| Statistical methods (Z-score / IQR) | Points deviating from the mean/quantiles are anomalies | Fast, interpretable, only suitable for univariate/simple distributions |
| Isolation Forest | Random partitioning; anomaly points are easily "isolated" (short path length) | Efficient, resistant to high dimensions, the go-to choice for tabular anomaly detection |
| LOF (Local Outlier Factor) | Compare a sample's density with its neighbors | Can detect local anomalies |
| One-Class SVM | Learn the boundary of "normal samples" | Effective in high dimensions, sensitive to hyperparameter tuning |
| Autoencoder reconstruction error | Normal samples have small reconstruction errors, anomalies have large ones | For deep learning scenarios |
Anomaly detection is fundamentally an "out-of-distribution" problem: normal data accounts for the vast majority, and the model only needs to describe "what normal looks like" — points that deviate too much are anomalies. This shares the same philosophy as [unsupervised clustering]'s "main cluster structure" concept.
Association Rules: Market Basket Analysis
Apriori (1994): discovers "if you buy A, you also often buy B" rules from transaction data. Three metrics:
- Support: the proportion of transactions covered by the rule (P(A∩B));
- Confidence: the conditional probability of B appearing given A appeared;
- Lift: P(B|A)/P(B), >1 indicates a positive correlation.
A classic case: Walmart's "beer and diapers" (actually a famous example from a data mining textbook). Applications: bundle sales, shelf placement, cross-recommendation. Today in large-scale scenarios, it has been replaced by collaborative filtering and embedding-based recommendation, but the interpretability of rules remains a practical tool for marketing campaign design (see Recommender Systems).
How Unsupervised Learning Lands in Practice
Unsupervised learning is rarely "delivered standalone"; it serves more as an upstream component of a pipeline:
Four Landing Patterns of Unsupervised Learning
├── Preprocessing: clustering → model per cluster (build separate churn models after customer segmentation)
├── Feature: dimensionality reduction/autoencoder → serve as features for downstream supervised models
├── Engine: representation pre-training → fine-tuning (both BERT and CLIP are unsupervised pre-trained)
└── Fallback: anomaly detection → data cleaning, fraud alerts (can go live without labels)This explains why unsupervised learning is "cheap but useful": it doesn't rely on expensive labeling, but its outputs (groups, representations, anomalies) can be directly consumed by supervised learning or business operations.
Unsupervised vs Self-supervised vs Weakly Supervised
- Unsupervised: uses only x, with no labels whatsoever;
- Self-supervised: constructs pseudo-labels from x itself (predict masked words in a sentence, predict image rotation angle) — essentially "creating supervised signals in an unsupervised manner," which is the core of BERT/GPT pre-training;
- Weakly supervised: uses imprecise/incomplete labels (remote supervision, rule-based labeling). The highlight moment for "unsupervised" in the era of large models is actually self-supervised — see Generative Models for details.
Tradeoffs
- K-Means vs DBSCAN: data is spherical and you know K → K-Means; arbitrary shapes, want automatic noise removal → DBSCAN.
- PCA vs t-SNE/UMAP: need interpretability, reversibility, variance preservation → PCA; only need visualization → UMAP/t-SNE.
- Clustering vs Dimensionality Reduction: clustering gives "groups," dimensionality reduction gives "coordinates"; the two are often combined (reduce dimensions first, then cluster, for speed/visualization).
- Automation vs Interpretability: the "correct answer" for unsupervised learning is defined by humans — every clustering result must go back to the business to validate its meaning, otherwise it's just "looks like it's grouped."
Further Reading
- Supervised Learning — the labeled counterpart paradigm
- Reinforcement Learning — the third paradigm
- Clustering and Dimensionality Reduction — breakdown of representative algorithms in practice
- Feature Engineering — dimensionality reduction as a feature processing step
- Generative Models — the unsupervised → self-supervised evolution
- Data and Data Engineering — anomaly detection in data cleaning
References
- scikit-learn: Clustering / Dimensionality reduction documentation
- MacQueen. Some methods for classification and analysis of multivariate observations (K-Means, 1967)
- Ester et al. A density-based algorithm for discovering clusters (DBSCAN, KDD 1996)
- Pearson. On lines and planes of closest fit (PCA, 1901)
- van der Maaten & Hinton. Visualizing Data using t-SNE (JMLR, 2008)
- Agrawal & Srikant. Fast Algorithms for Mining Association Rules (Apriori, VLDB 1994)
- Liu, Ting, Zhou. Isolation Forest (ICDM 2008)