Theme
Generative Models
Concept Definition: Learning to "Create" Rather than "Judge"
Discriminative models learn decision boundaries: given x, output class probability P(y|x) — classification and regression both fit here. Generative models learn the data distribution itself: estimating P(x) (or the joint distribution P(x,y)), so they can sample new samples from the distribution that are similar but different from the training data.
| Discriminative Model | Generative Model | |
|---|---|---|
| Learning objective | Decision boundary P(y | x) |
| Input/output | x → label | Noise/condition → new sample |
| Examples | Logistic regression, SVM, CNN classifier | GAN, VAE, diffusion models, GPT |
| In one line | "Is this a cat or a dog?" | "Draw a cat that doesn't exist" |
The essence of generative modeling: if the model truly learns the data distribution P(x), then sampling from P(x) produces convincing new data. The challenge is that distributions of high-dimensional data are extremely complex (an image of 1024×1024 is a manifold in a million-dimensional space), impossible to write explicitly, and can only be approximated.
Three Technical Routes
1. VAE: Variational Autoencoder (2013)
Idea: compress data into a low-dimensional latent space, then reconstruct from the latent space. The encoder maps x to a distribution of latent variable z (mean + variance), and the decoder reconstructs x from z. Trained via "variational inference": maximize "reconstruction likelihood - KL divergence" (keep the latent space close to a standard normal).
- Advantages: stable training, latent space is continuous and interpolatable (go halfway between z₁ and z₂ to generate an intermediate state);
- Disadvantage: generated images are blurry (Gaussian likelihood reconstruction leads to "mean blurring").
2. GAN: Generative Adversarial Network (2014)
Idea: two networks game against each other — the generator G creates fake samples (from noise), and the discriminator D distinguishes real from fake. G's goal is to fool D, D's goal is to catch G, adversarial training continues until G's samples are indistinguishable from real ones.
Generator G: noise z → fake sample G(z)
Discriminator D: sample → real/fake probability
Objective: min_G max_D E[log D(x)] + E[log(1 - D(G(z)))]- Advantages: generated samples are sharp and crisp (the 2014–2020 dominant force of image generation: face generation StyleGAN, image translation pix2pix);
- Disadvantage: unstable training (mode collapse: generates only a few types of samples), hard to converge — the "hyperparameter-tuning hell" reputation mainly comes from GANs.
3. Diffusion Models: Denoising Diffusion (2020)
Idea: first gradually add noise to data until it becomes pure noise (forward process), then learn a network that gradually denoises to restore the data (reverse process). Generation = start from pure noise, denoise step by step to get a clear image.
Forward (during training): x₀ → add noise → x₁ → … → x_T (pure noise)
Reverse (during generation): pure noise → denoise → … → x₀ (image)- Advantages: stable training, generation quality is the current best (since 2022, dominating text-to-image: Stable Diffusion, DALL·E, Midjourney, Imagen are all diffusion models);
- Disadvantage: slow generation (requires dozens to thousands of denoising iterations), needs accelerated sampling (DDIM, LCM) and distillation.
Route Comparison
| VAE | GAN | Diffusion | |
|---|---|---|---|
| Training stability | Stable | Poor (prone to collapse) | Stable |
| Generation quality | Blurry | Good but lacks diversity | Best |
| Diversity | Good | Poor (mode collapse) | Good |
| Sampling speed | Fast | Fast | Slow (needs acceleration) |
| Current status | Theoretical foundation | Still used in some domains | Standard for image/audio generation |
Conditional Generation: Making Generation Controllable
Unconditional generation is just "randomly draw an image"; business needs conditional generation — generate content matching a given prompt/label/image:
- Conditional diffusion (text-to-image): Stable Diffusion uses a text encoder (CLIP) to turn prompts into conditions, injected into the denoising network;
- ControlNet: use edge maps, poses, depth maps to control composition;
- Text-to-image workflow: prompt → diffusion sampling → upscale (super-resolution) → post-processing, see Diffusion Models and Generative AI.
Autoregressive Generation and Large Language Models
Another generation route is autoregressive: treating generation as a sequential "predict the next token one by one" problem — that's what GPT does:
P(x) = P(t₁)·P(t₂|t₁)·P(t₃|t₁,t₂)·…
Generation = sample the next token one by one, stitch into full textLarge language models (LLMs) are essentially conditional autoregressive generative models: given a prompt, generate a response token by token. Their abilities (translation, summarization, coding, reasoning) all come from the seemingly simple goal of "learning to predict the next word" — scaling laws cause "predicting the next word" to exhibit understanding and reasoning. For the full discussion, see Large Language Models (LLM) and Transformers and NLP.
Multimodal generation: text, images, audio, and video are converging — CLIP embeds images and text in the same space, and text-to-image, image-to-text, and text-to-video (Sora, etc.) all share the same "conditional generation" philosophy.
The Evaluation Challenge of Generative Models
There's no "answer key" for generative results, and evaluation is much harder than for discriminative models. Three categories of metrics, each with tradeoffs:
| Metric | Measures | Limitation |
|---|---|---|
| FID (images) | Feature distance between generated and real distributions | Needs many samples, biased toward statistics |
| IS (Inception Score) | Generation quality × diversity | Sensitive to "classes," unreliable for non-natural images |
| Human evaluation | Real user experience | Expensive, slow, but the ultimate standard |
| BLEU/ROUGE (text) | Overlap with reference text | "Right but different" generative text gets scored low |
Industry consensus: the final judge of generation quality is human (or downstream task performance). Industry commonly uses a three-layer system: "automated metrics for initial screening + human spot-checking + downstream task evaluation." Evaluation methodology is covered in Building an Evaluation System from Scratch.
Capabilities and Risks of Generative AI
Generative models bring new risks that discriminative models don't have, and practitioners must face them:
- Hallucination: generative models "nonsense with confidence" — inventing non-existent citations and facts. The essence is that they learn "sequences that look human," not "fact databases";
- Copyright and compliance: training data contains copyrighted material, and generated results may infringe (the compliance issues of text-to-image and code generation have been increasingly controversial since 2023);
- Deepfakes: the abuse of face-swap and voice-forging technology;
- Data pollution: AI-generated content flowing back into training sets, potentially degrading models ("model inbreeding") — this is a new challenge in data engineering post-2024;
- Alignment: models can be induced to generate harmful content, requiring safety alignment (RLHF, guardrails) — see Large Language Models (LLM).
Generation ≠ Fact
The goal of generative models is "distribution similarity," not "factual correctness." Using LLMs as search engines or treating generated images as real evidence will lead to pitfalls. The engineering positioning of generative AI is as a "creative tool," requiring verification and human judgment as fallback.
Tradeoffs
- Quality vs. speed: the more steps in a diffusion model, the sharper but slower — use distillation/accelerated sampling to trade for real-time performance;
- Controllability vs. diversity: the stronger the condition (ControlNet, low temperature sampling), the more controllable but less diverse, tuned by business per scenario;
- Self-build vs. use existing APIs: training generative models is extremely expensive (GPU clusters + massive data), so unless it's core business,prefer to use existing models (Stable Diffusion, GPT API) for fine-tuning/prompt engineering;
- Creativity vs. risk: generative AI is a double-edged sword; deployment must include content moderation and compliance evaluation.
Further Reading
- Deep Learning Fundamentals — the network foundation of generative models
- Diffusion Models and Generative AI — applied breakdown of text-to-image
- Large Language Models (LLM) — the most powerful form of autoregressive generation
- Transformers and NLP — architecture of generative NLP
- Interpretability and Fairness — the trustworthiness of generative models
- Data and Data Engineering — the data pollution problem of AI-generated content
References
- Kingma & Welling. Auto-Encoding Variational Bayes (VAE, ICLR 2014)
- Goodfellow et al. Generative Adversarial Nets (GAN, NeurIPS 2014)
- Ho, Jain, Abbeel. Denoising Diffusion Probabilistic Models (DDPM, 2020) — the foundational paper on diffusion models
- Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models (Stable Diffusion, 2022)
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021)
- Brown et al. Language Models are Few-Shot Learners (GPT-3, 2020)
- Heusel et al. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (FID, 2017)