Skip to content

Multimodal Models

Quick overview Multimodal models enable AI to understand text, images, audio, and video simultaneously. This article unpacks contrastive alignment in CLIP/ALIGN, the "visual encoder + projector + LLM" architecture of LLaVA/Qwen-VL/GPT-4V, early/late/cross-attention fusion schemes, Gemini/Omni unified models, MMMU evaluation, multimodal Agents and RAG, and current limitations.

Multimodal Models ​

In a sentence: Multimodal models process text, images, audio, video, and other modalities within a single model, aligning representations across modalities so they correspond to and reason about each other — it is the critical step for deep learning to move from "single-sense" to "general intelligence," and is the most active model family after Large Language Models (LLM).

1. Why Multimodal? ​

The limits of single-modality: text models "know" giraffes have long necks but haven't "seen" them; vision models have "seen" but can't explain. Multimodal benefits come in three layers:

  1. Alignment complements information: Images provide commonsense details to text; text provides abstract relationships to images;
  2. Mutual supervision for training: Image-text pairs, speech-text pairs are themselves free "pseudo-labels" — weak supervision signals at scales far exceeding manual annotation;
  3. Unified interface: One model handles "image captioning, speech understanding, Q&A, creation" — the foundation for interaction and Agents.

What supports this is the representation learning philosophy: different modalities, after encoding, become comparable and computable in a shared vector space (framework in Representation Learning and Pretraining).

2. Modality Alignment: CLIP and ALIGN ​

CLIP (2021, OpenAI) is a milestone in multimodal alignment: using 400 million (image, text) pairs for contrastive learning — pulling paired image-text vectors closer, pushing unpaired ones apart. After training, CLIP's image and text encoders share a semantic space: image-to-text retrieval, text-to-image retrieval, zero-shot classification ("assign this image to the closest class among {dog, cat, bird}") all become vector nearest-neighbor problems.

ALIGN (2021, Google) scaled data to 1.8 billion pairs, using "weakly aligned" web images with their alt-text for the same contrastive learning, demonstrating that data noise can be compensated by scale. The "dual-tower contrastive alignment" pattern established by CLIP/ALIGN is the common foundation for subsequent VLMs and text-to-image systems (Stable Diffusion's text encoder; see Diffusion Models and Generative AI).

3. Vision-Language Models: The VLM Stack ​

VLMs (Vision-Language Models) are the dominant form for "image understanding" today, with highly unified architecture — a three-piece stack:

  1. Visual encoder: Cut images into patches and feed to ViT (or use a CLIP-pretrained vision tower), outputting image tokens (see the ViT section in CNNs and Computer Vision);
  2. Projector: Project image token dimensions into the LLM's word vector space, letting the two modalities "speak the same language";
  3. Large language model: Concatenate projected image tokens into text sequences and continue autoregressive generation (architecture details in Transformer Architecture).

Representative models: LLaVA (2023) used "visual instruction data" for SFT, proving the cost-effectiveness of this path; Qwen-VL (2023) added dynamic-resolution input and visual localization capabilities; GPT-4V (2023) is the benchmark for closed-source flagships. VLM capabilities cover visual Q&A, OCR, chart understanding, image captioning, and referring expression grounding — training methods (contrastive pre-training → instruction fine-tuning → RLHF) mirror the three-step process of LLMs.

4. Text-to-Image / Text-to-Video: Cross-Modal Generation ​

Beyond alignment, another main thread is generating one modality from text:

  • Text-to-image: Diffusion models + CLIP text encoder = Stable Diffusion/DALL·E (principles in the "Diffusion Models and Generative AI" article);
  • Text-to-video: Sora (2024) and others tokenize video and apply spatiotemporal diffusion, a breakthrough for "generation duration and consistency" — see Frontier Advances;
  • Text-to-audio / audio-to-text: AudioLDM, Whisper, etc. — see Speech and Audio.

Cross-modal generation and understanding share the same foundational components (visual encoders, text encoders, diffusion/autoregressive generators) — this is exactly the footnote of multimodal "grand unification" trends.

5. Modality Fusion: Early, Late, and Cross-Attention ​

The core design question for multimodal models is "when and how to fuse":

Fusion MethodApproachCharacteristics
Early fusionConcatenate modality tokens at input (feed sequences together into a single Transformer)Most interaction, but requires high data volume and compute; Gemini falls in this category
Late fusionEncode each modality independently, then merge (concatenate/weighted/gated) for decision-makingReusable single-modality models, simple engineering; shallow information interaction
Cross-attentionOne modality as query, another as key-value (Encoder-Decoder style)Good for understanding tasks (VQA), controllable training

In practice, hybrid approaches are common: visual encoder (late component) → projector → fused with text via LLM self-attention (effectively "fusing at a deep position"). "Which layer to fuse at" has no universal answer — it depends on data, tasks, and compute budgets.

6. Unified Models: Gemini and Omni ​

"One model for all modalities and tasks" is the ultimate form:

  • Gemini (2023, Google): Native multimodal — from pretraining, text/images/audio/video are unified into token sequences fed to a single Transformer, rather than "adding a vision tower afterward"; long-context and cross-modal reasoning are its selling points;
  • Omni path (GPT-4o et al., 2024): Speech-to-speech direct mapping — no intermediate text conversion for audio, but direct cross-modal mapping, enabling low-latency, emotionally rich real-time voice conversations;
  • Any-to-Any: Input and output both support multiple modalities (images, audio, text), represented by Meta's ImageBind/AnyMAL and various unified large models.

The cost of unified models is extremely expensive training, and modalities may "mutually drag each other down" (catastrophic-forgetting-style modality degradation). The academic community is still exploring the optimal balance of "shared vs. specialized parameters."

7. Evaluation: MMMU and Multimodal Benchmarks ​

  • MMMU (2023): A multimodal multiple-choice benchmark spanning university-level disciplines (30 fields, cross-disciplinary reasoning) — the primary benchmark for VLM comprehensive ability;
  • Others: MMBench, MM-Vet (capability breakdown), OCRBench (document understanding), Video-MME (video);
  • Methodology and critique: A key difficulty in multimodal evaluation is separating "did it see?" from "did it reason?" (answering wrong could be an OCR failure, not a reasoning failure); image-text leakage into training sets and lagging benchmark updates are also common — evaluation methodology is covered in Deep Learning Evaluation and Experimentation and Evaluation in Practice.

8. Multimodal Agents and RAG ​

Multimodal models are developing "actuation capabilities":

  • Multimodal RAG: Retrieval databases include not just text, but also images, videos, audio — concatenate retrieved multimodal evidence into prompts for VLMs to answer, addressing knowledge timeliness and hallucinations — the multimodal extension of RAG in Large Language Models (LLM);
  • Multimodal Agents: VLMs understand screenshots and call tools (click, type, drag), enabling GUI Agents that "operate computers/phones by looking at screens" — a hot topic in Agent research (Agent fundamentals in Large Language Models (LLM));
  • Robotics / embodied AI: VLMs provide "scene understanding + instruction following," combined with Deep Reinforcement Learning Applications, the mainstream technical route for embodied intelligence.

9. Limitations and Open Problems ​

  • Hallucination amplification: VLM hallucinations on image details ("seeing" non-existent objects) are more subtle and harder to defend than in text models;
  • Alignment depth: CLIP-style alignment often stops at "surface semantics"; fine-grained spatial relationships, quantities, and temporal reasoning remain weak;
  • Data bottleneck: High-quality multimodal data is scarce, image-text pairs are noisy, and cultural differences (languages, scenes) cause general-purpose models to fail in long-tail scenarios;
  • Cost: Multimodal training and inference compute demands far exceed pure text, especially for long videos/high-resolution inputs;
  • Safety and fairness: Image generation and multimodal understanding both involve privacy, misinformation, and bias — governance discussions are in Interpretability and Fairness.

A Pragmatic Perspective

The gap between "high benchmark scores" and "product readiness" for multimodal models is bridged by substantial engineering: resolution, frame rate, latency, streaming, cross-language. Read papers for architecture understanding; for deployment, first get the "minimum viable loop" working before chasing fancy features.

10. Trade-offs ​

  • Native multimodal vs. external vision tower: Native (Gemini path) has a higher ceiling but is expensive; "pretrained visual encoder + projector + LLM" (LLaVA path) is flexible, reusable, and faster to prove value — most teams choose the latter;
  • Fusion position: Deep fusion has stronger interaction but harder training; shallow fusion is more stable but limited in effect — choose by task;
  • Unified vs. specialized: One model does everything vs. routing to modality-specialized models — the former is elegant but hard to optimize, the latter is engineering-mature but has fragmented interfaces;
  • Closed-source vs. open-source VLMs: Closed-source APIs are fast but constrained by data sovereignty; open-source (Qwen-VL, InternVL, LLaVA) is controllable but requires self-deployment and maintenance — trade-offs discussed in MLOps and Model Deployment.

Further Reading ​

References ​