Theme
Multimodal Models
In a sentence: Multimodal models process text, images, audio, video, and other modalities within a single model, aligning representations across modalities so they correspond to and reason about each other — it is the critical step for deep learning to move from "single-sense" to "general intelligence," and is the most active model family after Large Language Models (LLM).
1. Why Multimodal?
The limits of single-modality: text models "know" giraffes have long necks but haven't "seen" them; vision models have "seen" but can't explain. Multimodal benefits come in three layers:
- Alignment complements information: Images provide commonsense details to text; text provides abstract relationships to images;
- Mutual supervision for training: Image-text pairs, speech-text pairs are themselves free "pseudo-labels" — weak supervision signals at scales far exceeding manual annotation;
- Unified interface: One model handles "image captioning, speech understanding, Q&A, creation" — the foundation for interaction and Agents.
What supports this is the representation learning philosophy: different modalities, after encoding, become comparable and computable in a shared vector space (framework in Representation Learning and Pretraining).
2. Modality Alignment: CLIP and ALIGN
CLIP (2021, OpenAI) is a milestone in multimodal alignment: using 400 million (image, text) pairs for contrastive learning — pulling paired image-text vectors closer, pushing unpaired ones apart. After training, CLIP's image and text encoders share a semantic space: image-to-text retrieval, text-to-image retrieval, zero-shot classification ("assign this image to the closest class among {dog, cat, bird}") all become vector nearest-neighbor problems.
ALIGN (2021, Google) scaled data to 1.8 billion pairs, using "weakly aligned" web images with their alt-text for the same contrastive learning, demonstrating that data noise can be compensated by scale. The "dual-tower contrastive alignment" pattern established by CLIP/ALIGN is the common foundation for subsequent VLMs and text-to-image systems (Stable Diffusion's text encoder; see Diffusion Models and Generative AI).
3. Vision-Language Models: The VLM Stack
VLMs (Vision-Language Models) are the dominant form for "image understanding" today, with highly unified architecture — a three-piece stack:
- Visual encoder: Cut images into patches and feed to ViT (or use a CLIP-pretrained vision tower), outputting image tokens (see the ViT section in CNNs and Computer Vision);
- Projector: Project image token dimensions into the LLM's word vector space, letting the two modalities "speak the same language";
- Large language model: Concatenate projected image tokens into text sequences and continue autoregressive generation (architecture details in Transformer Architecture).
Representative models: LLaVA (2023) used "visual instruction data" for SFT, proving the cost-effectiveness of this path; Qwen-VL (2023) added dynamic-resolution input and visual localization capabilities; GPT-4V (2023) is the benchmark for closed-source flagships. VLM capabilities cover visual Q&A, OCR, chart understanding, image captioning, and referring expression grounding — training methods (contrastive pre-training → instruction fine-tuning → RLHF) mirror the three-step process of LLMs.
4. Text-to-Image / Text-to-Video: Cross-Modal Generation
Beyond alignment, another main thread is generating one modality from text:
- Text-to-image: Diffusion models + CLIP text encoder = Stable Diffusion/DALL·E (principles in the "Diffusion Models and Generative AI" article);
- Text-to-video: Sora (2024) and others tokenize video and apply spatiotemporal diffusion, a breakthrough for "generation duration and consistency" — see Frontier Advances;
- Text-to-audio / audio-to-text: AudioLDM, Whisper, etc. — see Speech and Audio.
Cross-modal generation and understanding share the same foundational components (visual encoders, text encoders, diffusion/autoregressive generators) — this is exactly the footnote of multimodal "grand unification" trends.
5. Modality Fusion: Early, Late, and Cross-Attention
The core design question for multimodal models is "when and how to fuse":
| Fusion Method | Approach | Characteristics |
|---|---|---|
| Early fusion | Concatenate modality tokens at input (feed sequences together into a single Transformer) | Most interaction, but requires high data volume and compute; Gemini falls in this category |
| Late fusion | Encode each modality independently, then merge (concatenate/weighted/gated) for decision-making | Reusable single-modality models, simple engineering; shallow information interaction |
| Cross-attention | One modality as query, another as key-value (Encoder-Decoder style) | Good for understanding tasks (VQA), controllable training |
In practice, hybrid approaches are common: visual encoder (late component) → projector → fused with text via LLM self-attention (effectively "fusing at a deep position"). "Which layer to fuse at" has no universal answer — it depends on data, tasks, and compute budgets.
6. Unified Models: Gemini and Omni
"One model for all modalities and tasks" is the ultimate form:
- Gemini (2023, Google): Native multimodal — from pretraining, text/images/audio/video are unified into token sequences fed to a single Transformer, rather than "adding a vision tower afterward"; long-context and cross-modal reasoning are its selling points;
- Omni path (GPT-4o et al., 2024): Speech-to-speech direct mapping — no intermediate text conversion for audio, but direct cross-modal mapping, enabling low-latency, emotionally rich real-time voice conversations;
- Any-to-Any: Input and output both support multiple modalities (images, audio, text), represented by Meta's ImageBind/AnyMAL and various unified large models.
The cost of unified models is extremely expensive training, and modalities may "mutually drag each other down" (catastrophic-forgetting-style modality degradation). The academic community is still exploring the optimal balance of "shared vs. specialized parameters."
7. Evaluation: MMMU and Multimodal Benchmarks
- MMMU (2023): A multimodal multiple-choice benchmark spanning university-level disciplines (30 fields, cross-disciplinary reasoning) — the primary benchmark for VLM comprehensive ability;
- Others: MMBench, MM-Vet (capability breakdown), OCRBench (document understanding), Video-MME (video);
- Methodology and critique: A key difficulty in multimodal evaluation is separating "did it see?" from "did it reason?" (answering wrong could be an OCR failure, not a reasoning failure); image-text leakage into training sets and lagging benchmark updates are also common — evaluation methodology is covered in Deep Learning Evaluation and Experimentation and Evaluation in Practice.
8. Multimodal Agents and RAG
Multimodal models are developing "actuation capabilities":
- Multimodal RAG: Retrieval databases include not just text, but also images, videos, audio — concatenate retrieved multimodal evidence into prompts for VLMs to answer, addressing knowledge timeliness and hallucinations — the multimodal extension of RAG in Large Language Models (LLM);
- Multimodal Agents: VLMs understand screenshots and call tools (click, type, drag), enabling GUI Agents that "operate computers/phones by looking at screens" — a hot topic in Agent research (Agent fundamentals in Large Language Models (LLM));
- Robotics / embodied AI: VLMs provide "scene understanding + instruction following," combined with Deep Reinforcement Learning Applications, the mainstream technical route for embodied intelligence.
9. Limitations and Open Problems
- Hallucination amplification: VLM hallucinations on image details ("seeing" non-existent objects) are more subtle and harder to defend than in text models;
- Alignment depth: CLIP-style alignment often stops at "surface semantics"; fine-grained spatial relationships, quantities, and temporal reasoning remain weak;
- Data bottleneck: High-quality multimodal data is scarce, image-text pairs are noisy, and cultural differences (languages, scenes) cause general-purpose models to fail in long-tail scenarios;
- Cost: Multimodal training and inference compute demands far exceed pure text, especially for long videos/high-resolution inputs;
- Safety and fairness: Image generation and multimodal understanding both involve privacy, misinformation, and bias — governance discussions are in Interpretability and Fairness.
A Pragmatic Perspective
The gap between "high benchmark scores" and "product readiness" for multimodal models is bridged by substantial engineering: resolution, frame rate, latency, streaming, cross-language. Read papers for architecture understanding; for deployment, first get the "minimum viable loop" working before chasing fancy features.
10. Trade-offs
- Native multimodal vs. external vision tower: Native (Gemini path) has a higher ceiling but is expensive; "pretrained visual encoder + projector + LLM" (LLaVA path) is flexible, reusable, and faster to prove value — most teams choose the latter;
- Fusion position: Deep fusion has stronger interaction but harder training; shallow fusion is more stable but limited in effect — choose by task;
- Unified vs. specialized: One model does everything vs. routing to modality-specialized models — the former is elegant but hard to optimize, the latter is engineering-mature but has fragmented interfaces;
- Closed-source vs. open-source VLMs: Closed-source APIs are fast but constrained by data sovereignty; open-source (Qwen-VL, InternVL, LLaVA) is controllable but requires self-deployment and maintenance — trade-offs discussed in MLOps and Model Deployment.
Further Reading
- Large Language Models (LLM) — The text brain and alignment paradigm of VLMs
- Transformer Architecture — The unified engine for cross-modal attention
- Diffusion Models and Generative AI — The generation engine for text-to-image/video
- CNNs and Computer Vision — The origin of visual encoders
- Speech and Audio — The addition of the audio modality
- Deep Learning Evaluation and Experimentation — Critical reading of MMMU and other benchmarks
References
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP) (ICML 2021)
- Jia et al. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision (ALIGN) (ICML 2021)
- Liu et al. Visual Instruction Tuning (LLaVA) (NeurIPS 2023)
- Bai et al. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond (2023)
- OpenAI. GPT-4V(ision) System Card (2023)
- Team Gemini. Gemini: A Family of Highly Capable Multimodal Models (2023)
- Yue et al. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI (CVPR 2024)
- Brooks et al. Video Generation Models as World Simulators (Sora) (2024)