Appearance
Multimodal Models
What Is a Multimodal Model?
A multimodal model is an AI model that can understand and generate information across multiple modalities — text, images, audio, video — such as GPT-4o, Gemini, and Qwen-VL. It is not "a model with a bolt-on image reader": from architecture to training data, multiple modalities are first-class citizens. One parameter set holds language's grammar and semantics alongside the pixel structure and object relations of images and the rhythm and emotion of sound.
Over the past decade, AI advanced along two parallel tracks: NLP rode the Transformer toward large language models, computer vision rode convolutions and ViT toward large image models, and speech stayed siloed in separate ASR/TTS systems. Multimodal models stitched these tracks together: a single input pipeline swallows text, images, recordings, and video, and the output spans modalities too. For ordinary users, the most visible change is that "AI can see and hear now" — upload a photo of a menu and ask it to split the bill per person; send a voice message and get a voice reply.
Today's mainstream chat products (ChatGPT, Gemini, Claude, and friends — see ChatGPT and Conversational AI) have almost all become multimodal entry points. Understanding multimodal models means understanding the default shape of the next generation of AI systems.
Why Multimodality Matters
Real-World Information Is Natively Multimodal
Humans have never perceived the world through text alone: in most situations, language is just a compressed encoding of reality, while images, video, and sound are the raw form of information. A surgical scan, footage from a construction site, a recording of a baby crying — this information either cannot be put into words or costs far too much to describe. A text-only model is born seeing the world through a keyhole.
"Omni-Modality" Is One Path Toward AGI
There is an industry consensus (best stated carefully): general intelligence requires an "all-modality" understanding of the world. A key source of human intelligence is integrating visual, auditory, linguistic, and tactile signals into one mental model in the same moment — a child learning the word "apple" sees its shape, feels its texture, and tastes it all at once. Correspondingly, exposing models to multiple modalities and having them learn cross-modal transfer is considered one candidate route toward AGI. That said, whether "omni-modality is a sufficient condition for AGI" remains unsettled — more modalities do not automatically mean stronger reasoning, and this is just one hypothesis among many on the frontier (see What Are AI's Hot Concepts).
More Natural Human-Computer Interaction
Typing is a legacy of the keyboard era; multimodality returns interaction to human instinct: speak, point at an image, wave a video. Real-time voice conversation cuts latency from "type, then read" to something close to talking with a person (see Speech AI: Whisper and TTS), and asking about an image beats typing a three-hundred-word description of it. This is not a UX polish — it is the step that turns AI from a "typewriter assistant" into a colleague at your side.
| Dimension | Text-only LLM | Multimodal model |
|---|---|---|
| Input | Text tokens | Text + images + audio + video |
| Output | Text | Text + images + audio (video) |
| Information density | Depends on how well the user describes things in words | Takes in raw signals, losslessly |
| Barrier to use | Keyboard | Natural language + the senses |
| World modeling | The world as described in words | The world as sensed |
The one-line verdict
Whether a model counts as "multimodal" is not about whether it can generate images — it is about whether its input crosses modalities and whether the modalities are jointly modeled in one parameter set. A pure text-to-image model (say, a standalone diffusion generator) usually does not qualify as a full multimodal model.
Core Technical Routes
Modality Encoding: Turning Everything into Vectors
Before any modality enters the model, it is chopped into "digital fragments" and mapped to vectors. Each modality "tokenizes" very differently:
| Modality | Encoder | Unit | Representative approach |
|---|---|---|---|
| Text | Tokenizer (BPE, etc.) | Subword tokens | GPT, Llama families |
| Image | ViT vision encoder | Fixed-size patches (e.g. 16×16) | CLIP-ViT, SigLIP |
| Audio | Spectrogram / sample encoders | Spectrogram frames or discrete tokens | Whisper encoder, HuBERT |
| Video | Per-frame image encoding + temporal modeling | Frame sequences / 3D patches | ViViT, Video-LLaMA |
Text keeps the Transformer family's tokenizer. On the image side, the landmark is ViT (Vision Transformer): slice an image into patches, flatten each patch into a vector, and feed them to a Transformer as "image tokens" — ViT is essentially a Transformer with the tokens swapped from words to image patches. This is why one architecture unified CV and NLP: the same attention mechanism models word tokens and patch tokens alike.
Audio has two routes: the spectrogram route converts waveforms into mel spectrograms (treated like images), and the discrete-token route quantizes audio into acoustic tokens (SoundStream-style codecs, common in speech LLMs). Either way, it ends up in the same Transformer framework.
Alignment and Fusion: Making Modalities Understand Each Other
Encoding is followed by the hard part — putting the representations of different modalities into one shared semantic space. History offers three generations of mainstream solutions:
Generation 1: Two-Tower Contrastive Learning (CLIP)
OpenAI's 2021 CLIP (Contrastive Language-Image Pre-training) established the paradigm for image-text alignment: two towers — an image encoder and a text encoder — each encode "the image" and "its caption" into vectors, and contrastive learning pulls correctly paired vectors together while pushing mismatched ones apart. Trained on 400 million image-text pairs scraped from the internet, the result is that image vectors and text vectors share a semantic space — "a photo of a dog" and "dog" sit close together.
Generation 2: Vision Encoder + Projection Layer (LLaVA)
2023's LLaVA grafted the CLIP idea onto an LLM: freeze a CLIP vision tower and a large language model, connect them with a trainable MLP projection layer that maps image features into the text embedding space, then fine-tune on high-quality instruction data generated by GPT-4 — producing an open-source assistant that can "look at a picture and talk about it." The Q-Former (introduced by BLIP-2) is an alternative projector: a lightweight Transformer "queries" image features and compresses them into the few vectors most relevant to text, saving parameters. This generation became known as vision-language models (VLMs).
LLaVA-style architecture (alignment route 2.0):
Image ──→ CLIP-ViT ──→ patch features ──┐
├──→ MLP projection ──→ [image tokens] ─┐
Text ──→ tokenizer ──→ [text tokens] ──────────────────────────────────────────┴─→ LLM (autoregressive generation)Generation 3: A Unified Token Space (GPT-4o)
2024's GPT-4o went further: all modalities are turned into tokens and enter a single neural network. Text no longer passes through a "projection" step; images and audio are likewise chopped into tokens, mixed with text tokens, and attended to together — cutting input-output latency to an average of 320 milliseconds, close to the rhythm of human conversation. Google's Gemini instead takes the "natively multimodal" route: joint training on text, images, audio, and video from the very start of pretraining, rather than "train a text LLM first, then bolt on vision."
| Route | Representative | Strength | Cost |
|---|---|---|---|
| Two-tower contrastive learning | CLIP (2021) | Training data is easy to get; strong retrieval/alignment | Aligns but doesn't generate; needs a separate decoder |
| Projection bridge | LLaVA, BLIP-2 (2023) | Reuses mature LLMs; open-source and reproducible | The projection layer loses information; weaker cross-modal reasoning |
| Unified token space | GPT-4o, Gemini (2024) | End-to-end, low latency, balanced abilities | Extreme demands on training data and compute |
The architectural thread
From CLIP to LLaVA to GPT-4o, the essence is deepening alignment: CLIP aligns two encoders; LLaVA pushes visual features "into" the language model; GPT-4o simply lets all modalities natively share one set of parameters. Grasp this thread and you grasp the decade-long design arc of multimodal models (read alongside Anatomy of the Overall Architecture).
The Generation Side: Understanding, Reversed
"Generation" in multimodality belongs to different technology stacks depending on the modality:
| Generation task | Core technology | Related reading |
|---|---|---|
| Text-to-image | Diffusion models | Diffusion Models and Generative AI, Midjourney and Image Generation |
| Text-to-video | Diffusion + temporal modeling (DiT) | Sora and Video Generation |
| Speech synthesis (TTS) | Acoustic model + vocoder / discrete-token generation | Speech AI: Whisper and TTS |
| Mixed image-text output | Joint token autoregression / diffusion branches | GPT-4o, Gemini, and other flagships |
One point that is easy to confuse: many "multimodal models" only do understanding (a VLM reads images and outputs text), while generation is handled by a separate diffusion model. Only true "omni-modal" flagships close the loop of "image in, image/voice out." When evaluating products, check understanding and generation on separate spec sheets.
Key Models Compared
Multimodal models from different routes differ sharply in capability boundaries. Check this table before choosing:
| Model | Image understanding | Video understanding | Audio input | Speech output | Image/video generation | Open source | Positioning |
|---|---|---|---|---|---|---|---|
| GPT-4o | ✅ | ✅ | ✅ | ✅ (native) | ⚠️ via DALL·E / Sora | ❌ | Omni-modal conversation flagship |
| Gemini 2.x | ✅ | ✅ (long videos) | ✅ | ✅ (native) | ✅ (Gemini image / Veo) | ⚠️ partial | Natively multimodal, ultra-long context |
| Claude 3.x | ✅ | ⚠️ limited | ❌ (as of 3.7) | ❌ | ❌ | ❌ | Vision understanding + long documents |
| Qwen-VL / Qwen2.5-VL | ✅ | ✅ | ✅ (Qwen family) | ✅ (Qwen family) | ⚠️ within the family | ✅ | The open-source multimodal benchmark |
| LLaVA | ✅ | ❌ | ❌ | ❌ | ❌ | ✅ | Vision-language assistant (research) |
| InternVL | ✅ | ⚠️ partial | ❌ | ❌ | ❌ | ✅ | Top-tier open-source VLM performance |
Currency check
The table reflects public information as of mid-2025 (dataAsOf: 2025-06), and model iteration is extremely fast: later Claude versions may already have added audio, and the Qwen family ships new releases monthly. Always defer to the vendor's latest documentation before choosing (see Models and Leaderboards Quick Reference).
Three trends fall out of this table:
- Understanding modalities are becoming standard: image understanding is now nearly a baseline for large models, and video understanding is trickling down from flagships;
- Speech is the dividing line: models with native voice input and output remain rare, because they demand low-latency end-to-end pipelines;
- Generation is mostly kept separate: most vendors let "the understanding LLM" and "the diffusion generator" each do their own job, then stitch them together at the product layer.
Evaluation and Challenges
Multimodal Hallucination
Multimodal hallucination is text hallucination's multimodal twin: the model describes objects that aren't there with a straight face, calls the male singer in the photo a female singer, or "reads" wrong numbers off a chart. Because images are information-dense and hard to verify pixel by pixel, multimodal hallucination is stealthier and more dangerous than the text kind — in medical imaging or financial charts, a single hallucination can cause real losses. Mitigations include dedicated adversarial evaluation sets, fine-tuning on fine-grained referring data, and teaching the model to say "I can't see that clearly" (see LLM Evaluation and Benchmarks).
How to Evaluate Cross-Modal Alignment
Multimodal evaluation goes far beyond "answer questions about a picture." The widely accepted evaluation dimensions today:
| Dimension | Representative benchmarks (as of 2025) | What it measures |
|---|---|---|
| Visual question answering | MMMU, MM-Vet, MathVista | Whether the model answers correctly when the information is in the image |
| Fine-grained perception | POPE, HallusionBench | Hallucination, misattribution, missed details |
| Cross-modal reasoning | MMMU, MMBench | Multi-step reasoning over charts and diagrams |
| Long-video understanding | Video-MME, LongVideoBench | Extracting information from minute-long videos |
| Spoken dialogue | SpeechArena, seed-bench (speech) | Naturalness of voice interaction, accent robustness |
| Instruction following | MMMT, MultiModalBench | Multi-turn instructions in multimodal settings |
The one-line verdict
When evaluating a multimodal model, test hallucination first, then reasoning, then instruction following. A model with a sky-high MMMU score but explosive POPE hallucination rates is usually unusable in real business.
Training Data: Scale and Quality
Multimodal is demanding on data in both quality and quantity: image-text pair datasets (LAION-5B, DataComp) are extremely noisy, and mismatched pairs ("the image is a dog, the caption says cat") are everywhere; video annotation costs several times more than for static images; and speech data raises speaker privacy and copyright issues. Data mixing ratios (the weights of text : image : audio : video) are among the most mystical yet most consequential engineering decisions in flagship model training — get the mix wrong and the model "underperforms by subject." See the Datasets and Tools Archive.
Application Scenarios
| Scenario | What it does | Typical forms |
|---|---|---|
| Image Q&A | Snap a photo and ask; summarize what's in an image | Asking about a recipe from a photo, debugging from a screenshot |
| Document parsing | Layout understanding and extraction for PDFs, scans, receipts | Contract review, invoice OCR + semantics |
| Video understanding | Long-video search, highlight localization, safety monitoring | Sports highlights, security alerts |
| Embodied AI | A robot's "eyes" — sense the environment, then act | Robotic-arm grasping, robot vacuum navigation |
| Real-time voice conversation | Low-latency speech in, speech out | Simultaneous interpretation, customer service, in-car assistants |
| Multimodal creation | Text-to-image, image-to-image, video generation | Design drafts, short videos, e-commerce assets |
Embodied AI deserves its own mention: the core problem for robots and autonomous driving is not "generation" but "acting on perception," and multimodal vision-language models happen to supply the middle layer of "see → understand → turn into instructions" — one of the hottest foundations in robotics today (for the full discussion of robot agents, see AI Agents).
How Multimodality Relates to LLMs, Diffusion Models, and Agents
- Multimodality is the LLM's sensory extension: a language model is a "thinker without senses"; multimodality gives it eyes, ears, and a voice. GPT-4o's underlying reasoning descends directly from GPT-4, but only with vision and speech encoders attached does the "see + think + speak" loop close. So understanding multimodality starts with understanding Large Language Models (LLMs) and their reasoning mechanisms (reasoning models like DeepSeek-R1 are extending into multimodality too).
- Diffusion models split the generation job: diffusion excels at reconstructing pixels from noise (the low-level generation of images/video/audio), while multimodal LLMs excel at "understanding + planning + orchestration." Flagship products are typically "the LLM decides, the diffusion model executes." Each mechanism is covered in Diffusion Models.
- Multimodality is the agent's window onto the world: an agent that only reads text can only "read the world"; a multimodal agent can "see the world" — operate software from screenshots, plan a trip from traffic maps, understand users from video. Multimodal + agents is the most active product direction since 2024 (see Manus and Agent Apps); and for an agent to "remember" the images it has seen and retrieve multimodal content, it needs vector databases and semantic search.
- Fusion with RAG and knowledge graphs: in enterprise deployments, multimodality is often combined with Retrieval-Augmented Generation (RAG) into "multimodal RAG" (retrieve image/document fragments and feed them to the VLM), while on the knowledge side, image entities are merged with text entities in a knowledge graph.
- Knock-on effects on engineering: multimodal models are parameter-heavy with long inputs; production depends on inference optimization and quantization and deployment practice; adjustments for vertical businesses go through fine-tuning and PEFT; safety boundaries involve AI safety and governance.
The relationship in one diagram:
Diffusion models (pixel generation) ──┐
├─→ Multimodal models (unified understanding + planning) ──→ Agents (perceive → decide → act)
LLMs (language reasoning) ────────────┘Trade-Offs
- End-to-end unification vs. modular assembly: a unified token space gives low latency and balanced abilities but costs a fortune to train and is an engineering black box; a modular pipeline (encoder + projection + LLM + diffusion) is composable and swappable but loses information between modules. Small teams should start with the modular route, then evolve toward end-to-end.
- The "greedy modalities" trap: the more modalities, the harder the data mixing, the more complex evaluation, and the larger the hallucination surface. In real products, "image + text only" is often sturdier than "force-fed omni-modality."
- The gap between open and closed flagships is narrowing: Qwen-VL, InternVL, and the LLaVA line have pushed open-source VLMs to near-closed-source levels; but heavy modalities like speech and video still call for closed flagships or in-house work.
- Latency vs. capability: native speech at conversational latency is the dividing line for the "multimodal conversation" experience — the price is a complex inference chain (streaming + interruptions + prosody control). See Speech AI.
Further Reading
- Transformers and Attention — ViT is a Transformer too; the foundation of multimodality
- Large Language Models (LLMs) — the "thinking core" of multimodal models
- Diffusion Models and Generative AI — the underlying engine of image/video generation
- AI Agents — multimodality is the agent's window onto the world
- LLM Evaluation and Benchmarks — the full methodology for modality hallucination and cross-modal evaluation
- Midjourney and Image Generation — a text-to-image case study
- Sora and Video Generation — a text-to-video case study
- Speech AI: Whisper and TTS — the engineering picture for speech recognition and synthesis
- Manus and Agent Apps — what multimodal agents look like as products
- Models and Leaderboards Quick Reference — first-hand quick reference for each model's multimodal abilities
References
- Radford et al., Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021) — the founding work of image-text contrastive learning
- Liu et al., Visual Instruction Tuning (LLaVA, 2023) — visual instruction tuning; the de facto standard for open-source VLMs
- Li et al., BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs (2023) — where the Q-Former comes from
- OpenAI, GPT-4V(ision) System Card (2023) — the systematic account of GPT-4V's capabilities and risks
- OpenAI, Hello GPT-4o (2024) — the GPT-4o release notes and omni-modal conversation demos
- Gemini Team, Gemini: A Family of Highly Capable Multimodal Models (2023) — the natively multimodal route paper
- Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT, 2020) — the foundational Vision Transformer paper
- Yue et al., MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark (2023) — the college-level multimodal reasoning benchmark
- OpenAI, Sora technical report — the diffusion/DiT route for video generation