Theme
Multimodal Large Language Models
Multimodal LLMs are large models capable of simultaneously understanding and generating information across multiple modalities — text, images, audio, and video. They extend the "language intelligence" of LLMs to "perceiving the world": answering questions about images, conversing via voice, understanding videos, and even generating images and video. During 2023–2025, multimodal capability evolved from a "nice-to-have" to a standard feature of flagship models — GPT-4o, Gemini, Claude, and Qwen-VL all include it without exception. This page breaks down the two technical approaches and representative models; for the language-side foundation, see The GPT Series: From GPT-1 to GPT-4o.
I. What Is a Multimodal LLM: From "Pure Language" to "Full-Sensory"
Pure text LLMs have token sequences as both input and output; multimodal LLMs extend input/output to include images, audio, and video. Modality coverage can be divided into several tiers:
| Capability Tier | Examples | Description |
|---|---|---|
| Visual Understanding | GPT-4V, LLaVA, Qwen-VL | Image comprehension: identification, Q&A, chart reasoning, OCR |
| Visual + Language Generation | GPT-4o, Gemini | Can generate images alongside text, create chart-generating code |
| Full Modal Omni | GPT-4o, Gemini 2.x | Unified text + image + audio + video understanding and output |
| Video Generation | Sora, Veo | Text/image-to-video generation |
| World Model Exploration | Sora Technical Report | Video generation as a "physical world simulator" |
A commonly misunderstood point about "multimodal": it's not simply "adding an image input box." The real benefit of multimodal models is their ability to perform "cross-modal reasoning" — describing what you see, identifying objects by sound, converting actions in a video to text, and cross-verifying modalities (e.g., "whether the license plate in the image matches the textual description"). Multimodal systems with input but no cross-modal reasoning are "multichannel" rather than truly "multimodal." To evaluate a multimodal model, focus on "cross-modal consistency" tasks (whether images and text corroborate each other, whether audio aligns with subtitles) rather than just looking at each modality's standalone performance.
II. Two Technical Approaches: Modular vs. Native Multimodal
1. Modular (Adapter) Approach: Visual Encoder + Projector + LLM
Approach: "A visual encoder converts images into vectors, a projector maps those vectors into the LLM's token space, and the LLM generates the output."
Image → [Visual Encoder CLIP ViT] → Image Vectors → [Projector] → LLM Token Space
↓
Text → [tokenizer] → tokens ────────────────────────────────────→ [LLM] → Text Output- Advantages: Reuses mature text LLMs without retraining the base model; fast development and modular (easy to swap LLMs or visual encoders);
- Disadvantages: Limited information sharing between vision and language — image details (small objects, spatial relationships, temporal changes) can easily be lost; inference overhead requires loading the full visual backbone.
Representative: LLaVA, Qwen-VL, and the original GPT-4V approach, all using the "CLIP + projector + Vicuna/Qwen" architecture.
2. Native (Omni) Multimodal: Unified Token Space
Approach: Tokenize images, audio, and video directly, training end-to-end with text tokens in the same Transformer, with shared attention across modalities.
- Advantages: Better cross-modal alignment, supports true "multimodal reasoning" and smooth cross-modal interaction (describing images, identifying objects by sound);
- Disadvantages: Large training data and compute requirements, engineering complexity.
Representative: GPT-4o (one model handling text + vision + audio + video), Gemini 1.5/2.x.
| Comparison Dimension | Modular (Modular) | Native (Native) |
|---|---|---|
| Architecture | Encoder + projector + LLM | Unified token space, end-to-end |
| Development Cost | Low (reuse text LLM) | High (retrain from scratch) |
| Visual Detail Preservation | Moderate | Better |
| Cross-Modal Reasoning | Limited | Strong |
| Representative Models | LLaVA, Qwen-VL, early GPT-4V | GPT-4o, Gemini |
The two approaches are not mutually exclusive
Real-world models are often "hybrids": Qwen-VL uses a visual encoder but trains deeply enough; GPT-4o is called native multimodal but internal details are not public. The criterion should be "depth of modality alignment" rather than marketing language.
By 2025, the "modular vs. native" debate has faded — because the two approaches have converged in practice. Even models classified as native multimodal may retain visual encoder modules internally; and modular models, through deeper training, can achieve alignment quality approaching native. What truly distinguishes them is: whether modalities share attention and training signals, and whether cross-modal reasoning requires external components. When selecting, focus on "multimodal evaluation scores + actual business outcomes" rather than marketing claims.
III. Representative Model Evolution: A Multimodal Chronicle
| Time | Model | Key Point |
|---|---|---|
| 2021.01 | CLIP | Image-text contrastive learning, 400M pairs, de facto visual encoder standard |
| 2022.04 | Flamingo | Frozen visual encoder + frozen LLM, Perceiver resampler + gated cross-attention |
| 2023.04 | LLaVA | Simplest modular approach: linear projector + visual instruction fine-tuning, open-source hit |
| 2023.08 | Qwen-VL | Chinese multimodal open-source; Qwen2-VL (2024.8) brought significant improvements |
| 2023.09 | GPT-4V | Vision version of GPT-4 via open API, multimodal enters closed-source flagships |
| 2023.12 | Gemini 1.0 | Google's native multimodal approach launched |
| 2024.02 | Gemini 1.5 Pro | Million-token context, supports long video understanding |
| 2024.05 | GPT-4o | Full-modal real-time voice conversations, multimodal goes free |
| 2024.02/12 | Sora | Video generation "world simulator," from preview to full release |
| 2025 | GPT-5, Gemini 2.x, Claude 4 | Full multimodal + reasoning + Agent integration |
1. CLIP: The "Visual Foundation Component" of Multimodal
CLIP (2021) used 400 million "image-text" pairs for contrastive learning: pulling matching image-text pairs closer, pushing non-matching ones apart, producing aligned image-text representations. It became the visual encoder for virtually all modular multimodal models, making "text-to-image retrieval and image-to-text retrieval" a general capability.
2. Flaminging: Elegant Modular Assembly of Frozen Models
DeepMind's Flamingo (2022) proved that you don't need to modify the weights of LLMs or visual encoders — simply adding a "resampler (Perceiver Resampler) + gated cross-attention" in the middle allows frozen LLMs to learn image-description skills with very few new parameters.
3. LLaVA: The "Simplest Route" for Open-Source
LLaVA (2023) made the approach extremely minimal: CLIP ViT + linear projector + Vicuna (a dialogue model based on Llama), then fine-tuned visual-language instruction following using GPT-4-generated instruction data (see below). It proved the effectiveness of "modular + instruction tuning" at the lowest cost, becoming the foundational hit of open-source multimodal models.
4. Qwen-VL: The Chinese Representative of Open-Source Multimodal
Alibaba's Qwen-VL series (2023.8 → Qwen2-VL 2024.8 → Qwen2.5-VL 2025.1) excels at document understanding, OCR, video understanding, and Chinese-language scenarios. Combined with the Llama and open-source ecosystem weight openness strategy, it has become one of the top choices for enterprise on-premises multimodal deployment.
5. GPT-4V / GPT-4o: Two Leaps by Closed-Source Flagships
- GPT-4V (2023.9): Added visual input to GPT-4, enabling it to read charts, analyze screenshots, recognize handwriting, and reason;
- GPT-4o (2024.5): Full-modal real-time voice (average response ~320ms), one model that simultaneously "sees, hears, and speaks," bringing multimodal into everyday life for the masses.
It's important to recognize the capability boundaries of multimodal models: they excel at "description and understanding" but still fall significantly short of humans in "precise spatial reasoning" (geometry problems, mapping), "temporal causality" (cause-effect in video events), and "fine-grained quantity comparison" (how much bigger one object is than another). Any claim that "multimodal models surpass humans" needs to be verified with benchmark datasets. For high-risk applications (medical imaging, autonomous driving), multimodal models can only serve as "assisted screening," with final judgments requiring human review.
6. Gemini: Native Multimodal + Ultra-Long Context
Google's Gemini has followed a native multimodal approach from 1.0 (2023.12); Gemini 1.5 Pro (2024.2) pushed context to approximately 1 million tokens, enabling understanding of entire videos and books; Gemini 2.x further integrated multimodal + Agent. Multimodal and long context (see Context and Long Context) converge on Gemini.
7. Sora: Video Generation and the "World Simulator"
Sora (2024) is OpenAI's video generation model, capable of producing high-fidelity videos from text or images. Its significance goes beyond "text-to-video": OpenAI positions it as a "world simulator" for understanding the dynamics of the physical world — if video generation learns "how objects move and how light and shadow change," this itself becomes one path to modeling the world (this assertion is exploratory; refer to official releases for the latest).
Looking back at the chronology, two structural inflection points emerge. The first was 2021–2022's "representation alignment": CLIP and Flamingo proved that "visual encoder + language model" could acquire visual-language capability at low cost, making multimodal go from "impossible" to "feasible." The second was 2023–2024's "full-modal native": GPT-4o and Gemini brought audio and video into a unified token space, upgrading interaction from "answer about images" to "real-time conversation." The next candidate inflection point is "multimodal reasoning and generation closed-loop": models not only understand images but also generate and self-verify them (Sora-like world simulation), and integrate multimodal perception into Agent execution loops (see LLM-Based Agents). A common thread at each inflection point is that "alignment and coordination between modalities" advances further.
Following this chronology, a practical conclusion emerges: for closed-source, look at GPT-4o / Gemini; for open-source, look at Qwen-VL / LLaVA; for research, look at CLIP / Flamingo lineages. Closed-source provides the strongest experience and managed convenience; open-source enables privatization and customization; research lineages provide mechanistic interpretability. The three are not in competition but in division of labor — when selecting, evaluate all three dimensions: capability, cost, and controllability, not just leaderboard scores.
IV. Visual-Language Instruction Tuning: Teaching Models to "Read Images on Command"
Like text models, multimodal models need alignment to be "usable." Visual Instruction Tuning is a key step in this process:
| Stage | Data | Purpose |
|---|---|---|
| Pretraining Alignment | Large-scale image-text pairs | Learn image-text representation alignment |
| Visual Instruction Tuning | Images + instructions + answers (e.g., LLaVA uses ~158K GPT-4-generated instruction data) | Learn to "look at images and answer per instructions" |
| Preference / Safety Alignment | Multimodal preference pairs, red-team samples | Usefulness, honesty, safety (see Alignment: RLHF and DPO) |
Data generation pattern: Use strong text models (GPT-4) to generate region descriptions, reasoning questions, and format constraints for images, then feed them to smaller models for fine-tuning — "using the strong to teach the weak" through data distillation is a common approach for multimodal models to grow rapidly.
Multimodal annotation is far more expensive than text annotation (image descriptions, region annotations require professional annotators), making synthetic data and distillation critical. Using strong models to generate image descriptions, programmatically generating chart Q&A pairs, and using simulation to generate video clips are all key techniques. Data quality can impact multimodal models even more than model scale, making "data engineering" a higher-priority function in multimodal teams than in pure text teams. Dataset resources and benchmarks are at Datasets and Benchmarks Archive.
V. Multimodal Evaluation: What Makes It Harder Than Pure Text
| Benchmark | What It Tests | Difficulty |
|---|---|---|
| MMMU | University-level multi-discipline image Q&A | Requires domain reasoning + visual detail |
| MMBench / Seed-Bench | Multimodal capability checklist | Broad coverage, fine-grained |
| MM-Vet | Comprehensive visual understanding | Emphasizes real tasks |
| POPE | Object hallucination detection | Tests whether models invent non-existent objects |
| BLINK / MathVista | Perception + math reasoning | Challenges the "reasoning from images" upper limit |
Multimodal evaluation has many metrics, but what matters most in practice is a "passing threshold" mindset: set a clear passing threshold for each key capability rather than pursuing overall scores. For example: "layout structure restoration rate" for document understanding, "hallucination rate upper bound" for visual Q&A, "temporal order accuracy" for video understanding. These thresholds come from business requirements (e.g., customer service requiring hallucination rate below a certain threshold), not from paper SOTA results. This approach aligns with the golden set method in Evaluations in Practice: the most important thing for multimodal projects is first defining "what counts as passing," then discussing "what scores are achieved." Having many metrics but no idea of the passing line is a common cause of multimodal project failures.
Two persistent challenges in multimodal evaluation: hallucination (models "seeing" objects that don't exist) and insufficient image-text association (large images with small targets, incorrect spatial relationships). Multimodal evaluation methodology and tools are covered in Evaluation and Benchmarks.
Multimodal models are not "search engines that can see images"
They see "pixel distributions," not "semantic databases." Complex charts, handwriting, and small-object identification still have significant failure rates; in high-risk scenarios like medical imaging and autonomous driving, human review is essential.
VI. Technical Details of Multimodal Models
1. Visual Encoding: From Pixels to Tokens
Whether modular or native multimodal, images must first be converted into "token sequences" the model can process:
| Step | Method | Effect |
|---|---|---|
| Patching | Divide image into 14×14/16×16 patches | Each patch becomes a token slot |
| Linear Embedding | Map each patch to a vector | Visual token |
| Position Encoding | Record the spatial position of each patch | Preserve spatial structure |
| Multi-Resolution | Higher-resolution images with more patches (e.g., Qwen-VL's resolution enhancement) | Support document/fine-grained recognition |
Note: the number of image tokens explodes with resolution (a 1024×1024 image has ~4096 patches), which is the root cause of "visual overhead" being far higher than text. Visual encoding is fundamentally the application of the attention mechanism (from Transformer Architecture Explained) to images.
2. CLIP: The Core of Contrastive Learning
| Element | Detail |
|---|---|
| Data | 400M image-text pairs |
| Objective | High scores for matching image-text pairs, low for non-matching (InfoNCE contrastive loss) |
| Technique | Large batches (tens of thousands of pairs) to ensure rich negatives |
| Output | Aligned visual and text encoders |
| Zero-Shot Capability | Classification via text prompts ("a photo of a cat") |
CLIP visual encoder + projector + LLM forms the skeleton of modular models like LLaVA.
3. Multimodal Training Data and Pipeline
| Stage | Data | Purpose |
|---|---|---|
| Contrastive Pretraining | Massive image-text pairs | Image-text representation alignment |
| Visual-Language Pretraining | Images + descriptions | Model learns to "talk about images" |
| Instruction Tuning | Images + instructions + answers (GPT-4 distilled) | Answer per instructions |
| Preference Alignment | Multimodal preference pairs | Safety and usefulness (see Alignment: RLHF and DPO) |
Data is the lifeline of modality alignment: image-text pairs have noise, instruction data is expensive, and video data is especially scarce. Dataset resources are at Datasets and Benchmarks Archive.
4. Tokenizing Audio and Video
- Audio: Usually first transcribed via speech recognition, or directly tokenized via audio tokenizers (e.g., EnCodec-style approaches) that cut audio into discrete tokens;
- Video: Extract frames for visual encoding, then add temporal attention; long videos rely on context compression (see Context and Long Context);
- Unified Token Space: GPT-4o and Gemini place text/image/audio/video tokens into the same vocabulary and attention — this is the engineering essence of "native multimodal."
5. Why Multimodal Is Harder Than Pure Text
The difficulty of multimodal training and inference isn't just "more modalities" — there are several structural challenges. First, data misalignment — image-text pairs and video-text pairs naturally contain noise (images and descriptions don't perfectly match), and cross-modal alignment signals are far weaker than internal text sequence signals. Second, uneven information density — an image can contain as much information as hundreds of words in token form, but the model treats every token equally, diluting visual details. Third, evaluation is hard — "is an answer correct" is relatively objective for text, but "did the model understand the image" is hard to automatically judge, blurring the boundary between hallucination and perception errors. These challenges explain why multimodal capability curves lag behind pure text models, and why the "CLIP-style contrastive learning + instruction tuning" combination remains the dominant recipe. Evaluation challenges and corresponding methods are at Evaluation and Benchmarks.
The key to multimodal isn't "more," it's "alignment"
Modular approaches first align image-text representations; native approaches align all modalities in a unified token space. The more modalities, the harder the alignment and the more data required; multimodal without alignment is merely "multiple input channels," not "multimodal intelligence."
VII. Boundaries: Division of Labor with the "Multimodal Handbook"
This handbook (the Large Model Handbook) and another set in this series, the Multimodal Handbook, have the following boundary:
| Topic | Where |
|---|---|
| Multimodal LLMs: LLM-centric, text + image/audio/video understanding and generation | This page (Large Model Handbook, case-study perspective) |
| Visual-language alignment mechanisms, instruction tuning, multimodal evaluation | This page + related concept pages in the Large Model Handbook |
| General multimodal representation learning (contrastive learning, cross-modal embeddings), low-level modeling of images/audio/video | Multimodal Handbook |
| Generative model foundations: Diffusion models, VAEs, GANs, and text-to-image/text-to-video mechanisms | Multimodal Handbook |
| Speech recognition/synthesis, classic audio/video signal processing | Multimodal Handbook |
In one sentence: this page tells the story and technical selection of "LLMs growing eyes and ears"; the Multimodal Handbook explains "how the eyes and ears themselves work." Read both volumes together for multimodal application selection.
VIII. Future Trends in Multimodal
- Full-modal Omni becomes standard: After GPT-4o and Gemini 2.x, "one model handling text/image/audio/video simultaneously" becomes the flagship baseline;
- Multimodal + Agents: Vision agents (seeing screens, operating interfaces), audio agents (real-time meetings, customer service) integrate multimodal into task closed-loops (see LLM-Based Agents);
- Video generation moves toward "controllable generation": From "generating an acceptable video" to "precisely controlling content and physical laws";
- Multimodal evaluation standardization: Hallucination rate, spatiotemporal consistency, instruction following will become industry entry barriers.
A dimension often overlooked for multimodal's future: training and inference cost. Visual tokens are dozens of times more than text tokens, and audio/video even more, making full-modal models far more expensive to use than pure text. The industry is reducing costs from three directions: more compact visual encoding (fewer tokens expressing more information), modality distillation (teaching small models with large models), and "on-demand modality activation" (only processing modalities that actually appear in input). Understanding these cost structures is a prerequisite for multimodal selection.
Timeline advice for practitioners: at this stage, focus on mastering the "multimodal understanding + RAG + Agent" combination (document intelligence, visual Q&A), which is the most mature deployment direction. "Multimodal generation" (images/video) is mainly for content creation scenarios, where ROI depends on your business model. As for "world model" research, observe cautiously and invest sparingedly. Layering by maturity is a practical mindset for avoiding multimodal selection pitfalls.
One-sentence summary of the multimodal landscape: understanding is mature, generation is catching up, world models are the distant future — arrange your investment pace with this mental model.
One sentence to remember multimodal
Multimodal LLM = Language Brain + Sensory Organs: Modular approaches first let models "see" (CLIP + projector + LLM), native multimodal then lets models "understand and communicate" (GPT-4o, Gemini); the next step is making models "take action" (Agents).
IX. Multimodal Applications and Deployment Practices
1. Typical Application Scenarios
| Scenario | Input → Output | Representative Models |
|---|---|---|
| Document Understanding | Scanned/PDF → Structured Data | GPT-4o, Qwen2.5-VL |
| Screenshot Q&A | Error Screenshot → Solutions | GPT-4o, Gemini |
| Image Retrieval | Natural Language → Images | CLIP-family embeddings |
| Real-Time Voice Assistant | Voice → Voice | GPT-4o |
| Long Video Understanding | Video → Summary/Q&A | Gemini 1.5+ |
| Vision Agent | Screen → Actions | Operator-type (see LLM-Based Agents) |
| Text-to-Image/Video | Text → Images/Video | Sora, Veo, etc. |
2. Document Understanding: The Hottest Multimodal Deployment Scenario
Enterprise knowledge sits in PDFs, scanned documents, and tables in large volumes. Multimodal models replace "parse documents" with "read documents":
- Advantages: Read layouts directly (tables, formulas, seals), skip the OCR pipeline;
- Engineering Essentials: High-resolution patching (divide and reassemble documents), combine with RAG (see RAG: Retrieval-Augmented Generation);
- Risks: Complex layouts can still be misread; manual sampling required.
3. Multimodal Model Selection Table
| Need | Recommendation | Rationale |
|---|---|---|
| Strongest visual reasoning | GPT-4o / Gemini | Complex charts, long videos |
| Chinese document understanding | Qwen2.5-VL | Strong Chinese OCR/layout |
| Open-source privatization | Qwen-VL / LLaVA | Open weights |
| Real-time voice | GPT-4o | Low latency, full-modal |
| Image retrieval embedding | CLIP / SigLIP | Image-text alignment |
| Video generation | Sora / Veo | Quality-leading |
4. Common Misconceptions
| Misconception | Correct Approach |
|---|---|
| Thinking multimodal "can understand any image" | Complex layouts/small targets still fail; sampling required |
| Using vision models for OCR | OCR is a sub-capability; structured extraction needs specialized design |
| Multimodal = images + text is enough | Audio/video alignment matters just as much |
| Only trusting leaderboards | Use custom multimodal evaluation sets (see Evaluation and Benchmarks) |
| Ignoring privacy | Images/audio contain extensive personal info; compliance first |
5. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| Difference between GPT-4o and GPT-4V? | 4o is native full-modal real-time; 4V is the early modular vision version |
| Which open-source multimodal to choose? | Qwen2.5-VL is among the strongest overall |
| How expensive are image tokens? | A high-res image has ~thousands of tokens; include in cost estimates |
| Which benchmarks for multimodal evaluation? | MMMU, MMBench, POPE (hallucination), and more |
| How to do video understanding? | Frame extraction + long context (Gemini approach) |
| How to combine with pure text models? | Multimodal handles perception, text models handle planning; often combined in Agents |
One sentence to remember multimodal deployment
First define "which modality solves which business problem," then select the approach (modular vs. native) and model. Multimodal cost and failure rates exceed pure text; use evaluation sets to uphold quality baselines.
X. Key Multimodal Papers and Evaluation Resources
1. Key Papers at a Glance
| Paper / Work | Year | One-Sentence Contribution |
|---|---|---|
| CLIP | 2021 | Image-text contrastive learning, visual encoder standard component |
| Flamingo | 2022 | Elegant modular assembly: frozen models + cross-attention |
| LLaVA | 2023 | Simplest modular + visual instruction tuning |
| InstructBLIP | 2023 | Instruction-tuned multimodal foundation model |
| Qwen-VL / Qwen2-VL | 2023–2024 | Chinese open-source multimodal benchmark |
| GPT-4V / GPT-4o | 2023–2024 | Closed-source flagship multimodal and full-modal |
| Gemini 1.0 / 1.5 | 2023–2024 | Native multimodal + million-token context |
| Sora | 2024 | Video generation as a world simulator |
2. Evaluation Resources Checklist
| Benchmark | Tests | Useful For Judging |
|---|---|---|
| MMMU | University multi-discipline visual Q&A | Domain reasoning ceiling |
| MMBench | Capability checklist items | Capability profile |
| MM-Vet | Comprehensive visual understanding | Overall capability |
| POPE | Object hallucination | Hallucination risk |
| SEED-Bench | Broad multimodal capability coverage | Selection reference |
| MathVista | Visual math reasoning | Charts and quantitative reasoning |
| BLINK | Human perception baseline | Fine-grained perception |
3. Evaluation Traps and Advice
- Visual input is hard to standardize: Different resolutions/crops of the same question yield very different results; evaluations need fixed input specifications;
- Hallucination is a hard issue: POPE-style evaluation should be included in selection criteria;
- Instruction bias: Models are sensitive to prompt styles; evaluation sets should cover real user phrasing;
- Cost: Visual evaluation has high token overhead; control batch sizes and retries.
One-sentence advice for selectors: multimodal leaderboard score differences within 2–3 points usually don't constitute a decision basis, because prompt styles, input specifications, and evaluation set versions all affect scores; what truly matters is performance on your own business data.
XI. Common Multimodal Engineering Pitfalls
| Pitfall | Symptom | Countermeasure |
|---|---|---|
| Resolution mismatch | Small text/fine-grained recognition fails | High-resolution patching + document-specific models |
| Messy input ordering | Multi-image relationship confusion | Fixed image numbering and ordering conventions |
| Misuse of caching | Cache invalidation fails on image hash changes | Cache keys should include image content hashes |
| Over-reliance on OCR | Treating layouts as plain text loses semantics | Structure preservation + layout understanding |
| Privacy compliance | Images/audio contain personal info | De-identification + compliance assessment first |
| Cost out of control | Large image tokens explode | Compression, patching, routing to smaller models |
5. Visual RAG: Combining Multimodal with Retrieval
A rapidly adopting deployment form is "Visual RAG": incorporating images, tables, and screenshots into the knowledge base, using multimodal embeddings (e.g., CLIP-family) for retrieval, then handing results to multimodal generation models for answering. Typical scenarios include: blueprint-centric knowledge bases ("where is this interface on the blueprint"), operations manuals with many screenshots, and audit systems requiring historical invoice comparisons. Engineeringly, note: image embeddings and text embeddings may not share the same space, requiring a unified "multimodal vectorization" approach; top-k recall rates for image retrieval are typically lower than for text, making the reranking step more critical. The difference between visual RAG and standard RAG is fundamentally a "modality alignment" issue — more details at RAG: Retrieval-Augmented Generation.
Multimodal deployment acceptance checklist
Run three types of tests before launch: ① capability testing (custom visual golden set); ② hallucination testing (POPE-style); ③ cost testing (tokens and latency per request). All three must pass before going live.
XII. Further Reading
- The GPT Series: From GPT-1 to GPT-4o — the language foundation for GPT-4V/4o
- Llama and the Open-Source Ecosystem — open-source multimodal lineages like Qwen-VL
- Alignment: RLHF and DPO — safety alignment for multimodal models
- Evaluation and Benchmarks — multimodal evaluation methodology
- Context and Long Context — the mechanism behind Gemini's million-token context
- LLM-Based Agents — the convergence of multimodal and Agents
- Frontier Progress — native multimodal and video generation frontiers
References
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021) — CLIP paper (arXiv)
- Alayrac et al. Flamingo: a Visual Language Model for Few-Shot Learning (2022) — Flamingo paper (arXiv)
- Liu et al. Visual Instruction Tuning (LLaVA, 2023) — LLaVA paper (arXiv)
- OpenAI. GPT-4V(ision) system card (2023) — GPT-4V technical report (arXiv)
- OpenAI. Hello GPT-4o (2024.5) — GPT-4o official release
- Google. Introducing Gemini: our largest and most capable AI model (2023.12) — Gemini 1.0 official blog
- Google. Introducing Gemini 1.5, our next-generation multimodal AI model (2024.2) — Gemini 1.5 official blog
- OpenAI. Sora: Creating video from text (2024) — Sora technical notes
- Yue et al. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark (2023) — MMMU evaluation benchmark paper (arXiv)