Theme
Diffusion Models and Generative AI
In a sentence: Diffusion models learn data distributions through a Markov chain of "progressive noise addition in the forward pass and progressive noise removal in the reverse pass" — first teach the model to "sculpt" pure noise into images step by step, then use that capability to generate new samples from random noise — combining the strengths of both VAEs and GANs (stable training, high quality, good coverage), they have become the mainstream generative engine for text-to-image, video, audio, and 3D generation since 2022.
1. Why Diffusion?
The fundamental question of generative modeling is "how to learn a sampleable, realistic data distribution." Let's look at the scars from two older approaches: VAE reconstruction loss pushes samples toward the "average," resulting in blurry images; GAN adversarial training produces sharp samples but is unstable in training and prone to mode collapse (see Generative Models for details). Diffusion models break through with a "boil-the-frog" learning approach — rather than predicting an entire image at once, it predicts "the small denoising step from the previous step" — each step is a simple task, and the network only needs to learn "how to denoise this noisy image a little." The overall effect: training is as stable as regression, yet generation is as sharp as adversarial training.
2. DDPM Principles
DDPM (Denoising Diffusion Probabilistic Models, Ho et al. 2020) defines two chains:
- Forward process (noise addition): Gradually add Gaussian noise to a clean image $x_0$ according to a fixed noise schedule $q(x_t | x_{t-1})$, approaching pure noise after $T=1000$ steps. Using reparameterization, you can compute any step $x_t = \sqrt{\bar\alpha_t},x_0 + \sqrt{1-\bar\alpha_t},\varepsilon$ in one shot;
- Reverse process (denoising): Train a neural network $\epsilon_\theta(x_t, t)$ to predict "the added noise $\varepsilon$." The training loss is extremely clean — it's just the mean squared error of the noise prediction: $$ \mathcal{L} = \mathbb{E}{t, x_0, \varepsilon}\left[|\varepsilon - \epsilon\theta(x_t, t)|^2\right] $$
- Sampling: Starting from pure noise, denoise iteratively using $\epsilon_\theta$ for $T$ steps to produce an image.
Intuitive analogy: break "generate an image" into 1000 steps of "wipe away a tiny bit of stain." The conditional distribution at each step is close to Gaussian and easy to learn, yet the chain as a whole captures complex real-world distributions. This "decompose complex tasks into simple ones" philosophy is analogous to "coarse-to-fine feature pyramids" in CNNs and Computer Vision.
Training Key Points
The training objective of DDPM is actually rooted in the same tradition as "denoising autoencoders" — predicting noise is equivalent to predicting the corrupted information. Using familiar backpropagation and automatic differentiation, you can train end-to-end without the two-network game that GANs require.
3. Sampling Acceleration: DDIM and DPM-Solver
DDPM requires 1000 steps for sampling, which is too slow. Two major acceleration approaches:
- DDIM (2021): Changes the reverse chain to be deterministic (no stochastic term for denoising), compressing steps to 20~50 while maintaining quality; also supports "step skipping" and unexpectedly provides "latent space interpolation" capability (smooth transitions between intermediate results of two samples);
- DPM-Solver (2022): Treats sampling as an ODE initial value problem, using high-order numerical solvers to compress steps to 10~20, the default sampler for many Stable Diffusion implementations.
Note the trade-off of sampling acceleration: too few steps loses detail and diversity, making it suitable for speed-sensitive scenarios (e.g., real-time image generation on the web).
4. Guidance: From "Random Generation" to "Precise Control"
Unconditional diffusion only generates random images. Guidance enables condition-based (text, class, image) generation:
- Classifier Guidance (2021): Uses gradients from an additional classifier to steer sampling toward "more like a given class" — good quality but requires training a separate classifier and taking gradient steps during sampling, making it slow;
- Classifier-Free Guidance (CFG) (2022): The same network learns both "conditional $\epsilon_\theta(x_t, c)$" and "unconditional $\epsilon_\theta(x_t, \varnothing)$," extrapolating at sampling time via $\tilde\epsilon = \epsilon_\theta(x_t,\varnothing) + s[\epsilon_\theta(x_t,c) - \epsilon_\theta(x_t,\varnothing)]$. Simply nullify the condition at a certain probability (e.g., 10%) during training — simple to implement, excellent results. CFG has become the standard for text-to-image. The larger the guidance strength $s$, the more aligned with the text, but at the cost of diversity (too strong introduces artifacts).
5. The Stable Diffusion Stack
Stable Diffusion (2022, Stability AI et al., based on Rombach et al.'s LDM) moves diffusion from "pixel space" to "latent space," enabling text-to-image generation with consumer-grade GPUs. Three components:
- VAE encoder/decoder: Compresses 512×512 images into 64×64×4 latent representations; diffusion operates only in latent space (saves computation while preserving details) — this is precisely the "VAE reborn" mentioned in the VAE and GAN section;
- UNet (diffusion backbone): A U-shaped network with residual connections and skip connections, internally using cross-attention to inject text conditions (the UNet skeleton originates from image segmentation; see the "CNNs and Computer Vision" article);
- CLIP text encoder: Encodes prompts into text condition vectors (CLIP itself is a contrastive-learning-pretrained image-text alignment model; see Multimodal Models).
LoRA fine-tuning (2021): Freezes the main UNet weights and trains only small low-rank decomposition matrices (rank r=8~64, reducing parameters by a thousandfold), fine-tuning on dozens of images to produce "a specific character/art style," with multiple LoRAs that can be stacked — it is the core tool of the community ecosystem and a textbook example of "efficient fine-tuning" from Representation Learning and Pretraining.
ControlNet (2023): Attaches a "trainable copy" beside the UNet, injecting condition structures like edges, poses, depth maps into generation, enabling "skeleton-controlled" output — solves the problem of pure text prompts having weak control over spatial structure, a milestone for controllable generation.
6. Application Landscape
- Text-to-image / Image-to-image: Midjourney, the Stable Diffusion ecosystem; image-to-image, inpainting, and outpainting are all mature features;
- Text-to-video: Sora (2024, OpenAI) tokenizes video and applies large-scale diffusion with spatiotemporal attention, a breakthrough for "generation duration and consistency"; open-source options include Video Diffusion series, CogVideoX, etc.;
- Audio: AudioLDM, Stable Audio, etc. do music and sound effect generation; diffusion-based TTS is also entering speech synthesis (see Speech and Audio);
- 3D: DreamFusion (2023) uses score distillation to let 2D diffusion guide 3D representation generation (NeRF/3DGS);
- Molecules / Materials: Diffusion models generate molecular conformations and crystal structures, a representative direction for scientific generation.
7. Evaluation and Limitations
- Evaluation metrics: FID (feature-level statistical distance between generated and real distributions, 2017) and CLIP score (semantic alignment between generated images and text) are the two major metrics for text-to-image; IS (Inception Score) was widely used early on. Metrics ≠ perceived quality: FID is sensitive to "distribution coverage and sharpness" but not to individual failures; for visual evaluation methodology, see Deep Learning Evaluation and Experimentation;
- Limitations: Slow sampling (an order of magnitude more expensive than GANs), occasional text-alignment failures (wrong finger counts, etc.), hallucinated details, high compute and memory costs, and copyright/ethics controversies around training data;
- Trends: Distillation (one-step generation, e.g., SD-Turbo/LCM), DiT (using Transformers to replace UNet, including Sora and SD3) — "diffusion + Transformer" is becoming the next-generation baseline. See Frontier Advances for the tech landscape.
Evaluation Caveats
A model that "looks good" doesn't mean it "understands semantics." High CLIP score doesn't mean logical correctness. For generative projects, always examine quantitative metrics, diversity, and failure-case analysis simultaneously, and build an human-evaluation dataset — see Evaluation in Practice for engineering evaluation practices.
8. Trade-offs
- Diffusion vs GAN vs VAE: Choose diffusion for quality and diversity; GANs or distilled diffusion still have value for real-time interaction and edge computing with tight compute budgets;
- Pixel space vs latent space: Latent space (LDM) saves compute but is resolution-limited by the VAE bottleneck; pixel-level (e.g., some high-res fine-tuning) preserves fidelity but is expensive;
- Speed vs quality: Sampling steps, CFG strength, and model size are all sliders — no free lunch;
- Controllability vs naturalness: Stronger ControlNet/LoRA means more stable control but may "overfit to prompts" and sacrifice generative diversity — treat trade-offs as product parameters, not as questions of technical correctness.
Further Reading
- VAEs and GANs — The two predecessors and component suppliers of diffusion models
- Generative Models — A unified perspective on probabilistic modeling
- Multimodal Models — CLIP text encoder and image-text alignment
- CNNs and Computer Vision — The visual backbone of UNet
- Transformer Architecture — The diffusion backbone of the DiT era
- Training Recipes and Hyperparameter Tuning — Hands-on experience with diffusion training
References
- Ho, Jain, Abbeel. Denoising Diffusion Probabilistic Models (DDPM) (NeurIPS 2020)
- Song, Meng, Ermon. Denoising Diffusion Implicit Models (DDIM) (ICLR 2021)
- Dhariwal, Nichol. Diffusion Models Beat GANs on Image Synthesis (Classifier Guidance) (NeurIPS 2021)
- Ho, Salimans. Classifier-Free Diffusion Guidance (NeurIPS 2022 Workshop)
- Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models (LDM/Stable Diffusion) (CVPR 2022)
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022)
- Zhang, Rao, Agrawala. Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet) (ICCV 2023)
- Heusel et al. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (FID) (NeurIPS 2017)