Skip to content

Generative Models

Quick overview From GANs to diffusion models to large language models — generative models are a class of methods that "learn to create." This article clarifies the essence of generative modeling (learning data distributions), the principles and comparisons of three major technical routes (VAE/GAN/diffusion), and the capabilities and risks of generative AI.

Generative Models ​

Concept Definition: Learning to "Create" Rather than "Judge" ​

Discriminative models learn decision boundaries: given x, output class probability P(y|x) — classification and regression both fit here. Generative models learn the data distribution itself: estimating P(x) (or the joint distribution P(x,y)), so they can sample new samples from the distribution that are similar but different from the training data.

Discriminative ModelGenerative Model
Learning objectiveDecision boundary P(yx)
Input/outputx → labelNoise/condition → new sample
ExamplesLogistic regression, SVM, CNN classifierGAN, VAE, diffusion models, GPT
In one line"Is this a cat or a dog?""Draw a cat that doesn't exist"

The essence of generative modeling: if the model truly learns the data distribution P(x), then sampling from P(x) produces convincing new data. The challenge is that distributions of high-dimensional data are extremely complex (an image of 1024×1024 is a manifold in a million-dimensional space), impossible to write explicitly, and can only be approximated.

Three Technical Routes ​

1. VAE: Variational Autoencoder (2013) ​

Idea: compress data into a low-dimensional latent space, then reconstruct from the latent space. The encoder maps x to a distribution of latent variable z (mean + variance), and the decoder reconstructs x from z. Trained via "variational inference": maximize "reconstruction likelihood - KL divergence" (keep the latent space close to a standard normal).

  • Advantages: stable training, latent space is continuous and interpolatable (go halfway between z₁ and z₂ to generate an intermediate state);
  • Disadvantage: generated images are blurry (Gaussian likelihood reconstruction leads to "mean blurring").

2. GAN: Generative Adversarial Network (2014) ​

Idea: two networks game against each other — the generator G creates fake samples (from noise), and the discriminator D distinguishes real from fake. G's goal is to fool D, D's goal is to catch G, adversarial training continues until G's samples are indistinguishable from real ones.

Generator G: noise z → fake sample G(z)
Discriminator D: sample → real/fake probability
Objective: min_G max_D  E[log D(x)] + E[log(1 - D(G(z)))]
  • Advantages: generated samples are sharp and crisp (the 2014–2020 dominant force of image generation: face generation StyleGAN, image translation pix2pix);
  • Disadvantage: unstable training (mode collapse: generates only a few types of samples), hard to converge — the "hyperparameter-tuning hell" reputation mainly comes from GANs.

3. Diffusion Models: Denoising Diffusion (2020) ​

Idea: first gradually add noise to data until it becomes pure noise (forward process), then learn a network that gradually denoises to restore the data (reverse process). Generation = start from pure noise, denoise step by step to get a clear image.

Forward (during training): x₀ → add noise → x₁ → … → x_T (pure noise)
Reverse (during generation): pure noise → denoise → … → x₀ (image)
  • Advantages: stable training, generation quality is the current best (since 2022, dominating text-to-image: Stable Diffusion, DALL·E, Midjourney, Imagen are all diffusion models);
  • Disadvantage: slow generation (requires dozens to thousands of denoising iterations), needs accelerated sampling (DDIM, LCM) and distillation.

Route Comparison ​

VAEGANDiffusion
Training stabilityStablePoor (prone to collapse)Stable
Generation qualityBlurryGood but lacks diversityBest
DiversityGoodPoor (mode collapse)Good
Sampling speedFastFastSlow (needs acceleration)
Current statusTheoretical foundationStill used in some domainsStandard for image/audio generation

Conditional Generation: Making Generation Controllable ​

Unconditional generation is just "randomly draw an image"; business needs conditional generation — generate content matching a given prompt/label/image:

  • Conditional diffusion (text-to-image): Stable Diffusion uses a text encoder (CLIP) to turn prompts into conditions, injected into the denoising network;
  • ControlNet: use edge maps, poses, depth maps to control composition;
  • Text-to-image workflow: prompt → diffusion sampling → upscale (super-resolution) → post-processing, see Diffusion Models and Generative AI.

Autoregressive Generation and Large Language Models ​

Another generation route is autoregressive: treating generation as a sequential "predict the next token one by one" problem — that's what GPT does:

P(x) = P(t₁)·P(t₂|t₁)·P(t₃|t₁,t₂)·…
Generation = sample the next token one by one, stitch into full text

Large language models (LLMs) are essentially conditional autoregressive generative models: given a prompt, generate a response token by token. Their abilities (translation, summarization, coding, reasoning) all come from the seemingly simple goal of "learning to predict the next word" — scaling laws cause "predicting the next word" to exhibit understanding and reasoning. For the full discussion, see Large Language Models (LLM) and Transformers and NLP.

Multimodal generation: text, images, audio, and video are converging — CLIP embeds images and text in the same space, and text-to-image, image-to-text, and text-to-video (Sora, etc.) all share the same "conditional generation" philosophy.

The Evaluation Challenge of Generative Models ​

There's no "answer key" for generative results, and evaluation is much harder than for discriminative models. Three categories of metrics, each with tradeoffs:

MetricMeasuresLimitation
FID (images)Feature distance between generated and real distributionsNeeds many samples, biased toward statistics
IS (Inception Score)Generation quality × diversitySensitive to "classes," unreliable for non-natural images
Human evaluationReal user experienceExpensive, slow, but the ultimate standard
BLEU/ROUGE (text)Overlap with reference text"Right but different" generative text gets scored low

Industry consensus: the final judge of generation quality is human (or downstream task performance). Industry commonly uses a three-layer system: "automated metrics for initial screening + human spot-checking + downstream task evaluation." Evaluation methodology is covered in Building an Evaluation System from Scratch.

Capabilities and Risks of Generative AI ​

Generative models bring new risks that discriminative models don't have, and practitioners must face them:

  • Hallucination: generative models "nonsense with confidence" — inventing non-existent citations and facts. The essence is that they learn "sequences that look human," not "fact databases";
  • Copyright and compliance: training data contains copyrighted material, and generated results may infringe (the compliance issues of text-to-image and code generation have been increasingly controversial since 2023);
  • Deepfakes: the abuse of face-swap and voice-forging technology;
  • Data pollution: AI-generated content flowing back into training sets, potentially degrading models ("model inbreeding") — this is a new challenge in data engineering post-2024;
  • Alignment: models can be induced to generate harmful content, requiring safety alignment (RLHF, guardrails) — see Large Language Models (LLM).

Generation ≠ Fact

The goal of generative models is "distribution similarity," not "factual correctness." Using LLMs as search engines or treating generated images as real evidence will lead to pitfalls. The engineering positioning of generative AI is as a "creative tool," requiring verification and human judgment as fallback.

Tradeoffs ​

  • Quality vs. speed: the more steps in a diffusion model, the sharper but slower — use distillation/accelerated sampling to trade for real-time performance;
  • Controllability vs. diversity: the stronger the condition (ControlNet, low temperature sampling), the more controllable but less diverse, tuned by business per scenario;
  • Self-build vs. use existing APIs: training generative models is extremely expensive (GPU clusters + massive data), so unless it's core business,prefer to use existing models (Stable Diffusion, GPT API) for fine-tuning/prompt engineering;
  • Creativity vs. risk: generative AI is a double-edged sword; deployment must include content moderation and compliance evaluation.

Further Reading ​

References ​