Skip to content

Generative Models

Quick overview Discriminative models answer "what is this"; generative models answer "generate one." This article surveys the spectrum of explicit/implicit density models, covering VAE (ELBO/reparameterization), GAN (adversarial game/mode collapse), diffusion models (DDPM), and autoregressive generation, their principles and relationships, and provides a selection guide.

Generative Models ​

One-sentence definition: A generative model learns the probability distribution of data and samples from it to produce new samples with the same distribution — discriminative models learn "P(label|data)," while generative models learn "P(data)" itself. This goal may seem more abstract, yet it has powered everything from image generation to language modeling, launching the entire era of generative AI. The crossover point in Representation Learning and Pretraining — "acquiring representations from generative tasks" — also lives here.

1. Generative Tasks: Sampling from Distributions ​

Given a dataset X = {x₁,…,x_N}, assume they come from some unknown distribution p_data(x). Generative modeling has two objectives:

  1. Fit: learn a parameterized p_θ(x) that approximates p_data(x).
  2. Sample: be able to draw new samples from p_θ(x) — new images, new text, new speech.

Why is "sampling from a distribution" harder than "classification"? Because the data space has extremely high dimensionality (a 512×512 image is 780,000 dimensions) and the distribution is extremely complex ("natural images" occupy only a tiny sliver of that vast space). The task of generative models is to "pinpoint" plausible points on this manifold.

Unconditional vs. conditional generation: unconditional generation samples from p(x); conditional generation samples from p(x|y) (given a class label or text prompt). The current mainstream (text-to-image, text-to-text) is almost entirely conditional generation.

2. Explicit vs. Implicit Density Models ​

The generative model family can be classified by "whether an explicit probability density is given":

  • Explicit density models: directly define a computable p_θ(x), trained via maximum likelihood (maximizing probability). Examples: autoregressive models, VAE (approximate density).
  • Implicit density models: don't write p_θ(x); only define a transformation that can sample from it (the generator). Example: GAN.
  • Diffusion models: somewhere in between — define a "stepwise denoising" process from noise to data, trained with an approximate density of "small Gaussians at each step." They are explicit but have likelihood that is hard to compute exactly.

This spectrum determines their training objectives and evaluation methods (whether NLL is computable, how to use FID); see the generative model evaluation section of Deep Learning Evaluation and Experiments.

3. VAE: Latent Variables, ELBO, and Reparameterization ​

Variational Autoencoder (VAE; Kingma & Welling, 2013) introduces a low-dimensional latent variable z, assuming x is generated by z: p_θ(x) = ∫p_θ(x|z)p(z)dz. This integral is intractable, so variational methods come to the rescue:

  • Use an encoder q_φ(z|x) to approximate the true posterior p(z|x).
  • The derivation of the optimization objective yields the ELBO (Evidence Lower Bound):
log p(x) ≥ E_{z~q}[log p(x|z)] − KL(q(z|x) ‖ p(z))

Each term has its role: the first is the reconstruction error (how well the decoder can restore z back to x); the second is a regularization term (keeping q close to the prior p(z), ensuring the latent space is continuous and samplable).

Reparameterization trick: sampling from q(z|x) is non-differentiable (no backpropagation possible), so we rewrite it as z = μ + σ⊙ε, ε~N(0,I) — "routing around" the randomness from the parameter path so gradients can flow through μ and σ (this is exactly the differentiability design in the computation graphs of Backpropagation and Automatic Differentiation).

VAE's positioning: likelihood is computable (NLL can be evaluated), the latent space has structure (enabling interpolation/manipulation), training is stable, but generated samples tend to be "blurry" (Gaussian decoder + MSE reconstruction error naturally produce blurriness).

4. GAN: Adversarial Game and Mode Collapse ​

Generative Adversarial Network (GAN; Goodfellow et al., 2014) introduces a game: generator G turns noise z into fake samples, discriminator D judges whether samples are real or fake. They are trained adversarially:

min_G max_D  E[log D(x)] + E[log(1 − D(G(z)))]
  • D's objective: distinguish real from fake (the max part).
  • G's objective: fool D (the min part).
  • The theoretical equilibrium: p_G = p_data, at which point D can only guess randomly (probability 0.5).

Two notorious chronic issues:

  • Training instability: both sides are "enemies" being updated dynamically. If D is too strong, G's gradients vanish; if G is too strong, D loses meaning. Engineering remedies: different learning rates (TTUR, two-time-scale updates, see Optimization and Gradient Descent), gradient penalty (WGAN-GP), spectral normalization.
  • Mode collapse: G learns only to generate a few "samples that fool D," losing all diversity. Remedies: improved losses, minibatch discrimination, diversity penalties.

GAN's positioning: high generation quality, fast inference (one-step generation), but difficult to train, no density, hard to converge to cover all modes. See VAE and GAN for details.

5. Diffusion Models: Forward Noising, Reverse Denoising, and DDPM ​

Diffusion models (ideas proposed by Sohl-Dickstein 2015; DDPM, Ho et al. 2020 made them practical) approach generation from another angle:

  • Forward process: progressively add Gaussian noise to data (x₀→x₁→…→x_T); after T steps, it approximates pure noise. The distribution of this process is known and computable in closed form.
  • Reverse process: learn a network to progressively denoise (x_T→…→x₀), predicting "the noise added at each step" (or the mean) at each step; the network output is the final sample.

DDPM's key tricks:

  1. Closed-form forward: x_t = √ᾱ_t·x₀ + √(1−ᾱ_t)·ε — the state at any step t can be computed in one step, no simulation needed.
  2. Extremely simple training objective: L = ‖ε − ε_θ(x_t, t)‖² — the network predicts the added noise, and MSE suffices.
  3. Stepwise sampling: start from x_T~N(0,I) and iteratively denoise for T steps (T is typically in the thousands; later methods like DDIM compressed sampling steps to dozens).

Diffusion models excel in both quality and diversity on images (surpassing GAN), powering text-to-image systems like Stable Diffusion; their combination with VAE (LDM diffusing in latent space) is standard engineering practice. The cost is slow sampling (iterative) and high inference cost. See Diffusion Models and Generative AI for the full picture.

6. Autoregressive Generation ​

Autoregressive generation decomposes sequence probability into conditional probabilities per token:

p(x₁,…,x_T) = Π_t p(x_t | x₁,…,x_{t−1})
  • Text: GPT series, "next token prediction," paired with causal masking from Attention. It is both a "language model" and a "generative model."
  • Images: PixelCNN/PixelRNN generate pixel by pixel; VQ-VAE discretizes images into tokens, then applies autoregression.
  • Audio: WaveNet generates sample point by sample point (see Speech and Audio).

Advantages of autoregression: likelihood is exactly computable (naturally supports NLL evaluation), training is stable; disadvantages: slow generation (depends on KV cache for acceleration, see the attention chapter) and long-range error accumulation.

7. Evaluation Metrics: NLL, FID, IS ​

"Good" generative models are hard to define in a single metric. The industry commonly uses three complementary metrics:

  • NLL (Negative Log-Likelihood): directly computable for explicit models (VAE, autoregressive); measures "fitting accuracy." Drawback: doesn't fully correlate with "human perception" (low NLL can still produce blurry images).
  • IS (Inception Score, 2016): generate samples, classify them with Inception-v3, and check "are they clear enough to classify (low class entropy) + are they diverse (high class distribution entropy)."
  • FID (Fréchet Inception Distance, 2017): the distance between two Gaussian distributions in the Inception feature space (mean + covariance) of real and generated images. More sensitive than IS at capturing blurriness and mode collapse; it is the current de facto standard for image generation.

Key reminder: these three metrics each measure different things. Use them in combination with human evaluation; see the evaluation chapter for definitions and pitfalls.

8. Three Generative Families: Comparison and Selection ​

FamilyTraining objectiveDensity computable?Training stable?Generation qualitySampling speedNotable examples
VAEVariational lower bound (ELBO)ApproximatelyStableMedium (blurry)Fast (one-step)VAE, VQ-VAE, LDM encoder
GANAdversarial gameNoUnstableHighFast (one-step)StyleGAN, BigGAN, GAN variants
DiffusionStepwise denoising (MSE noise prediction)ApproximatelyStableVery highSlow (iterative)DDPM, Stable Diffusion, Sora
AutoregressiveMaximum likelihood (per-token)ExactlyStableVery high (text)Slow (per-token)GPT, PixelCNN, WaveNet

Selection rules of thumb: image quality + speed → GAN; quality + diversity + accept slower sampling → diffusion; need latent semantics + likelihood → VAE; text/discrete sequences → autoregressive. Real-world products are often combinations: LDM = diffusion + VAE latent space, VQ-GAN = GAN + autoregressive tokenization.

Trade-offs

Likelihood computability vs. sample quality: autoregressive models with exactly computable NLL produce great text but are slower and more conservative for images; GAN never optimizes likelihood, yet its samples are strikingly impactful. "Fitting well" and "generating beautifully" are not the same thing.

One-step vs. iterative sampling: one-step (VAE/GAN) is fast but hard to cover complex distributions; iterative (diffusion) is steady but slow. DDIM/distillation are pushing iterative steps downward — the "fast vs. good" gap is being bridged by engineering.

Training stability vs. theoretical completeness: diffusion/autoregression have clean theory and stable training, at the cost of expensive sampling; GAN is cheapest at inference time, but training is like walking a tightrope. When selecting, consider both your inference budget (deployment costs in MLOps) and your training team's technical depth.

Generative models are spreading from images/text to video (Sora), audio, 3D, and scientific computing, but their spectrum — explicit/implicit, one-step/iterative, discrete/continuous — remains the map for understanding any new architecture. Together with Representation Learning and Pretraining and Loss Functions and Output Layers (the crossover of NLL/contrastive losses), they form a complete loop; one step further toward productization leads to Diffusion Models and Generative AI and Large Language Models (LLM).

Further Reading ​

References ​