Skip to content

Unsupervised Learning

Quick overview What to do without labels? Unsupervised learning discovers structure from data itself — clustering, dimensionality reduction, anomaly detection, and association rules. This article clarifies the principles of four major tasks, representative algorithms, evaluation challenges, and how "unsupervised" approaches land in real business.

Unsupervised Learning ​

Concept Definition: Learning without the Answer Key ​

Supervised learning learns from "labeled (x, y) pairs," while unsupervised learning deals with data where there's only input x, no label y. The goal isn't to predict an answer, but to discover the inherent structure of the data: which samples are similar? How many natural groups does the data have? Which dimensions matter? Which points are anomalous?

A one-line summary of the four major tasks:

  • Clustering: automatically group similar samples;
  • Dimensionality Reduction: compress high-dimensional data into lower dimensions while preserving the main structure;
  • Anomaly Detection: find points that differ significantly from the majority of samples;
  • Association: discover co-occurrence patterns between features/items.

The value of unsupervised learning is often underestimated in real business: in reality, labels are scarce and expensive (annotation costs money and people), while unsupervised learning can directly extract value from massive unlabeled data — customer segmentation, data compression, fraud detection, and feature learning. In the deep learning era, unsupervised "representation learning" (pre-training) has even become the core of large models — see Generative Models and Large Language Models.

Clustering: Grouping Data ​

Representative Algorithms ​

AlgorithmCore IdeaCharacteristics
K-MeansIterate: assign samples to nearest centroid → update centroidFast, simple, requires specifying K, sensitive to outliers/non-spherical distributions
DBSCANDensity-connected samples form clustersNo need to specify K, finds arbitrary shapes, auto-detects noise points
Agglomerative ClusteringBottom-up merging of nearest clusters, produces a dendrogramVisualization-friendly, view at any granularity
GMM (Gaussian Mixture)Assumes data is a mixture of K Gaussian distributionsSoft clustering (gives probabilities), can model elliptical clusters

The Full K-Means Pipeline ​

text
1. Select K initial centroids
2. Loop until convergence:
   a. Assign each sample to the nearest centroid (distance is typically Euclidean)
   b. Centroid of each cluster = mean of samples in that cluster
python
from sklearn.cluster import KMeans
import numpy as np

X = np.random.default_rng(0).normal(size=(500, 2))  # 500 two-dimensional samples
kmeans = KMeans(n_clusters=3, n_init=10, random_state=42)
labels = kmeans.fit_predict(X)
print(f"Centroids: {kmeans.cluster_centers_}")

How to Choose K? ​

  • Elbow method: plot the "intra-cluster sum of squares (inertia) vs. K" curve and look for the bend point;
  • Silhouette Score: measures "within-cluster compactness + between-cluster separation," pick the K with the highest score;
  • Business-driven: whether K=5 or K=8 for customer segmentation ultimately depends on whether the resulting groups are meaningful and interpretable in the business context.

The Evaluation Challenge of Clustering ​

Clustering has no "answer key," and evaluation is inherently difficult: internal metrics (silhouette score, inertia) only look at structure and don't guarantee business relevance; external metrics (ARI, NMI) require true labels — but when true labels exist, you'd typically just use supervised learning. Practical advice: prioritize business interpretability, use internal metrics only as a reference.

Dimensionality Reduction: Compressing Data while Preserving Structure ​

High-dimensional data has three problems: expensive computation, impossible visualization, and the curse of dimensionality (in high-dimensional space, samples become increasingly "sparse," making distance concepts meaningless). Dimensionality reduction maps data into a lower-dimensional space, divided into two categories:

Linear Dimensionality Reduction ​

PCA (Principal Component Analysis): finds orthogonal directions of maximum variance (principal components) and projects data onto them. Principle: the first principal component direction = the direction of maximum data variance. Uses of PCA:

  • Visualization (reduce to 2D/3D);
  • Denoising (discard components with low variance);
  • Feature compression (prevent overfitting);
  • Whitening/preprocessing (works with scaling).
python
from sklearn.decomposition import PCA

pca = PCA(n_components=2)      # reduce to 2 dimensions
X_2d = pca.fit_transform(X)
print(f"Explained variance ratio of two PCs: {pca.explained_variance_ratio_}")

Choosing the number of components: look at the "cumulative explained variance ratio" curve and pick the number that reaches 80%~95%.

Nonlinear Dimensionality Reduction ​

AlgorithmCore IdeaCharacteristics
t-SNEPreserves high-dimensional neighborhood relations, uses t-distribution for low-dim projectionGreat for visualization, but doesn't preserve global structure, results are stochastic and irreversible
UMAPManifold learning + topological ideasFaster than t-SNE, preserves global structure better, the go-to choice for visualization in recent years
AutoencoderNeural network: compress to a bottleneck then reconstructNonlinear representation learning, differentiable, can be fine-tuned

Three Misuses of Dimensionality Reduction

  1. Must standardize before PCA (otherwise features with large scales dominate the principal components);
  2. t-SNE distances are not trustworthy — it only preserves relative neighborhood relationships; "far" on the plot doesn't mean truly far;
  3. Dimensionality reduction for visualization ≠ for modeling: reducing X to 2 dimensions and then training a classifier usually loses information; use PCA/autoencoder, not t-SNE, for dimensionality reduction before modeling.

Anomaly Detection: Finding the "Odd Ones Out" ​

Business scenarios: fraud detection, industrial quality inspection, network intrusion, data cleaning.

MethodIdeaCharacteristics
Statistical methods (Z-score / IQR)Points deviating from the mean/quantiles are anomaliesFast, interpretable, only suitable for univariate/simple distributions
Isolation ForestRandom partitioning; anomaly points are easily "isolated" (short path length)Efficient, resistant to high dimensions, the go-to choice for tabular anomaly detection
LOF (Local Outlier Factor)Compare a sample's density with its neighborsCan detect local anomalies
One-Class SVMLearn the boundary of "normal samples"Effective in high dimensions, sensitive to hyperparameter tuning
Autoencoder reconstruction errorNormal samples have small reconstruction errors, anomalies have large onesFor deep learning scenarios

Anomaly detection is fundamentally an "out-of-distribution" problem: normal data accounts for the vast majority, and the model only needs to describe "what normal looks like" — points that deviate too much are anomalies. This shares the same philosophy as [unsupervised clustering]'s "main cluster structure" concept.

Association Rules: Market Basket Analysis ​

Apriori (1994): discovers "if you buy A, you also often buy B" rules from transaction data. Three metrics:

  • Support: the proportion of transactions covered by the rule (P(A∩B));
  • Confidence: the conditional probability of B appearing given A appeared;
  • Lift: P(B|A)/P(B), >1 indicates a positive correlation.

A classic case: Walmart's "beer and diapers" (actually a famous example from a data mining textbook). Applications: bundle sales, shelf placement, cross-recommendation. Today in large-scale scenarios, it has been replaced by collaborative filtering and embedding-based recommendation, but the interpretability of rules remains a practical tool for marketing campaign design (see Recommender Systems).

How Unsupervised Learning Lands in Practice ​

Unsupervised learning is rarely "delivered standalone"; it serves more as an upstream component of a pipeline:

Four Landing Patterns of Unsupervised Learning
├── Preprocessing: clustering → model per cluster (build separate churn models after customer segmentation)
├── Feature: dimensionality reduction/autoencoder → serve as features for downstream supervised models
├── Engine: representation pre-training → fine-tuning (both BERT and CLIP are unsupervised pre-trained)
└── Fallback: anomaly detection → data cleaning, fraud alerts (can go live without labels)

This explains why unsupervised learning is "cheap but useful": it doesn't rely on expensive labeling, but its outputs (groups, representations, anomalies) can be directly consumed by supervised learning or business operations.

Unsupervised vs Self-supervised vs Weakly Supervised

  • Unsupervised: uses only x, with no labels whatsoever;
  • Self-supervised: constructs pseudo-labels from x itself (predict masked words in a sentence, predict image rotation angle) — essentially "creating supervised signals in an unsupervised manner," which is the core of BERT/GPT pre-training;
  • Weakly supervised: uses imprecise/incomplete labels (remote supervision, rule-based labeling). The highlight moment for "unsupervised" in the era of large models is actually self-supervised — see Generative Models for details.

Tradeoffs ​

  • K-Means vs DBSCAN: data is spherical and you know K → K-Means; arbitrary shapes, want automatic noise removal → DBSCAN.
  • PCA vs t-SNE/UMAP: need interpretability, reversibility, variance preservation → PCA; only need visualization → UMAP/t-SNE.
  • Clustering vs Dimensionality Reduction: clustering gives "groups," dimensionality reduction gives "coordinates"; the two are often combined (reduce dimensions first, then cluster, for speed/visualization).
  • Automation vs Interpretability: the "correct answer" for unsupervised learning is defined by humans — every clustering result must go back to the business to validate its meaning, otherwise it's just "looks like it's grouped."

Further Reading ​

References ​