Theme
VAE and GAN
In one sentence: Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) are the two foundational paradigms of deep generative modeling—VAEs learn a sampleable latent space through variational inference, while GANs learn to forge data through a generative-adversarial game—together they defined the "AI creates content" era from 2014 to 2021, and they are the engineering protagonists behind the concept of Generative Models.
1. From Autoencoders to VAEs
Autoencoder (AE)
An autoencoder is a "copyist" that compresses input into a bottleneck and reconstructs it: encoder $x \to z$, decoder $z \to \hat x$, with reconstruction error $|x - \hat x|^2$ as the loss. The bottleneck forces the model to discard redundancy and retain key information—so $z$ becomes a form of representation. But AEs have a fatal flaw: their learned latent space is "fragmented"—any two random points may not be continuous, and randomly sampling z won't necessarily produce a valid x. So AEs can only do dimensionality reduction and denoising, not generation.
The VAE Breakthrough (2013, Kingma & Welling)
VAEs upgrade AEs into probabilistic models: instead of having the encoder output a single point for the latent variable, it outputs a distribution (mean μ and variance σ). This forces the latent space to be continuous and smooth: nearby z values generate nearby x values, and random sampling of z yields new samples.
Two core techniques:
- Reparameterization trick: Sample $z = \mu + \sigma \odot \varepsilon$ ($\varepsilon \sim \mathcal{N}(0, I)$). This turns an "unsampleable" operation into one that "comes from a fixed noise transformation"—so gradients can flow back through z to the encoder. This is the mathematical trick that makes end-to-end VAE training possible, and a textbook example of combining "probabilistic generative models + backpropagation" (see Backpropagation and Automatic Differentiation).
- ELBO objective: Maximize the evidence lower bound, equivalent to minimizing two terms: $$ \mathcal{L} = \underbrace{\mathbb{E}{z\sim q}[\log p(x|z)]}{\text{reconstruction}} - \underbrace{D_{KL}(q(z|x) | p(z))}_{\text{regularization}} $$ The regularization term pulls the encoder distribution toward a standard normal prior; the reconstruction term ensures decoding quality. The two terms pull in opposite directions: too much regularization → blurry reconstruction; too little → fragmented latent space. This tension runs through all subsequent VAE improvements (β-VAE, VQ-VAE).
2. GAN: The Game Between Generator and Discriminator
GANs (2014, Goodfellow et al.) take a completely different approach from VAEs: they don't explicitly model distributions; instead, they train a forger and an appraiser.
- Generator G: Takes noise z as input, outputs fake samples.
- Discriminator D: Judges whether an input is real or fake.
- The objective function is a minimax game: $$ \min_G \max_D ; \mathbb{E}{x\sim p{data}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1 - D(G(z)))] $$
Intuitive analogy: the forger constantly improves its craft, the appraiser constantly sharpens their eye, and they push each other to get stronger. The ideal equilibrium point is when the generated distribution $p_g = p_{data}$, at which point D outputs 0.5 everywhere—beaten. GAN's generated samples (especially images) are much sharper than VAEs because adversarial loss pushes for "realism" rather than "averaging," whereas VAE reconstruction loss (like L2) naturally converges to blurry averages. Note that training involves alternating updates of two roles—a game—unlike the single-objective gradient descent approach in Optimization and Gradient Descent.
3. Training Instability and Mode Collapse
GANs are notoriously hard to train, with two chronic issues:
- Training instability: The game has no unified convergence objective. If D is too strong, gradients vanish and G learns nothing; if D is too weak, G goes off course. Training is like "walking a tightrope," requiring careful balance of learning rates and update schedules for both sides (practical checklist in Training Recipes and Hyperparameter Tuning).
- Mode collapse: G discovers that "only producing a few types of realistic samples" is enough to fool D, so output diversity collapses (e.g., face generation only produces a few expressions). The essence is G's degenerate strategy: rather than cover the full distribution, cluster in a few spots.
Lessons from VAEs and GANs
VAEs are stable but blurry; GANs are sharp but collapse-prone—the trade-off between "training stability" and "sample quality" is the meta-proposition of generative model design. This contradiction is elegantly navigated in diffusion models and generative AI (both stable and sharp), but at the cost of slow sampling.
4. Improvement Lineage
| Model | Year | Solves | Core Idea |
|---|---|---|---|
| DCGAN | 2015 | Instability | All-convolutional architecture + design heuristics (no pooling, BatchNorm, LeakyReLU) |
| cGAN | 2014 | Controllable generation | Conditional input (class/text) injected into generator and discriminator |
| WGAN / WGAN-GP | 2017 | Collapse + instability | Replace JS divergence with Wasserstein distance, gradient penalty enforces 1-Lipschitz |
| CycleGAN | 2017 | Unpaired translation | Cycle consistency loss for style transfer |
| StyleGAN | 2018 | Controllability and quality | Disentangled latent space (style injection + mapping network), controllable face generation |
| BigGAN / StyleGAN2 | 2018–2020 | Large-scale high quality | Large batch + spectral normalization, peak of image generation |
The improvement thread is clear: smoother distance metrics (WGAN), conditional control (cGAN), disentanglement of latent dimensions (StyleGAN), and enhanced training stability (spectral normalization, gradient penalty). Conditional generation (cGAN) turns "generation" from random into "on demand," serving as the direct precursor to later text-to-image ("generate from text") models.
5. Applications
- Image editing and synthesis: Face editing (manipulating StyleGAN's latent space), super-resolution (SRGAN with perceptual + adversarial loss), inpainting, style transfer (CycleGAN)—all built on the visual features from CNN and Computer Vision.
- Domain translation: Day/night, seasons, sketch→photo—industrial applications include medical imaging cross-device enhancement.
- Anomaly detection: AE/VAE trained only on normal samples have high reconstruction error for anomalies, serving as a weakly supervised baseline for industrial quality control and financial risk.
- Data augmentation: Synthesize training data for underrepresented categories (but be cautious about distribution shift between synthetic and real data—evaluation methods in Deep Learning Evaluation and Experimentation).
6. Comparison and Current State with Diffusion Models
After 2020, diffusion models have comprehensively surpassed GANs in image generation (details in Diffusion Models and Generative AI):
| Dimension | VAE | GAN | Diffusion |
|---|---|---|---|
| Training stability | Good | Poor | Good |
| Sample quality / diversity | Blurry | Sharp but prone to collapse | Sharp and diverse |
| Sampling speed | Fast | Fast | Slow (requires accelerated sampling: DDIM, DPM-Solver) |
| Latent space controllability | Good (VAE) | Moderate | Moderate (guided by text/conditions) |
| Current state | Used as "latent compressor" within diffusion systems (VAE part of LDMs) | Stepped back in images; still valuable in video/point clouds | Mainstream for images/video/audio/3D |
A notable "revival of old ideas": today's diffusion systems often contain VAEs internally—the latent space of Stable Diffusion is obtained via VAE compression—and GAN's discriminator concept has inspired some distillation/acceleration methods. The three are not mutually exclusive but complementary components. See the "Generative Models" article for the full landscape.
7. Trade-offs
- VAE vs. GAN paradigm choice: Need stable training and good latent space (editing, interpolation, representation)? Choose VAE. Pursuing peak single-sample quality and willing to endure tuning pain? Choose GAN—but after 2024, most new scenarios default to diffusion.
- Quality vs. diversity: GANs naturally bias toward "safe high quality"; diffusion is more balanced.
- Controllability vs. freedom: Stronger conditional injection = more controllable, but also more easily "led astray by conditions."
- Compute cost: GAN sampling is extremely fast; diffusion sampling is expensive. For real-time interactive scenarios (e.g., virtual try-on), sticking with GAN can still be engineeringly rational—trade-offs in MLOps and Model Deployment.
Further Reading
- Diffusion Models and Generative AI—The next-generation paradigm taking over generation
- Generative Models—A unified perspective on probabilistic modeling
- CNN and Computer Vision—The visual foundation for generative image tasks
- Loss Functions and Output Layers—Adversarial loss, reconstruction loss, perceptual loss
- Optimization and Gradient Descent—The optimization essence of GAN's dual-optimizer game
- Training Recipes and Hyperparameter Tuning—Practical experience for stable GAN training
References
- Kingma, Welling. Auto-Encoding Variational Bayes (ICLR 2014)
- Goodfellow et al. Generative Adversarial Nets (NeurIPS 2014)
- Radford, Metz, Chintala. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks (DCGAN) (ICLR 2016)
- Arjovsky, Chintala, Bottou. Wasserstein GAN (ICML 2017)
- Gulrajani et al. Improved Training of Wasserstein GANs (WGAN-GP) (NeurIPS 2017)
- Karras, Laine, Aila. A Style-Based Generator Architecture for Generative Adversarial Networks (StyleGAN) (CVPR 2019)
- Mirza, Osindero. Conditional Generative Adversarial Nets (2014)
- Zhu et al. Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks (CycleGAN) (ICCV 2017)