Theme
Model Compendium
This page is a quick-reference guide to mainstream large language models: family, release date, parameter count, context window, open-source status, and highlights — along with a categorization by dimension, a licensing quick-reference, and a "how to choose a model" decision table. Model specs and version rollouts move fast — this page's data is current as of August 2025; always refer to official sources for the latest information. Closed-source model parameters are generally undisclosed and marked as "not disclosed."
Model profiles become outdated quickly
GPT-4o was once the leader, and now there's the o-series and GPT-5. Llama went from version 2 to version 4 in just two years. This page serves only as an "introductory map" — always check the official docs before procurement or model selection (see References at the end). For evaluation scores, adopt the critical perspective outlined in Evaluation and Benchmarks rather than fixating on leaderboards.
I. Model Overview Table
Representative versions by family. "Open source" means weights are publicly available (note that licenses vary by family). Parameter counts are either officially published or based on industry consensus; "not disclosed" where the information is unavailable.
| Family | Representative version | Release date | Parameters | Context | Open source? | Highlights |
|---|---|---|---|---|---|---|
| GPT (OpenAI) | GPT-3 | 2020-05 | 175B | 2K | No | First to demonstrate large-scale few-shot capability |
| GPT-4 | 2023-03 | Not disclosed (rumored MoE) | 8K / 32K | No | Exam-level capability, multimodal | |
| GPT-4o | 2024-05 | Not disclosed | 128K | No | Native multimodal, low-latency real-time interaction | |
| o1 | 2024-09 | Not disclosed | 200K | No | Test-time scaling (reasoning-time expansion) | |
| GPT-5 | 2025-08 | Not disclosed | ~400K | No | Native fusion of reasoning and tool use | |
| Claude (Anthropic) | Claude 3 series | 2024-03 | Not disclosed | 200K | No | Excellent long-context and safety policy |
| Claude 3.7 Sonnet | 2025-02 | Not disclosed | 200K | No | Hybrid reasoning (controllable thinking budget) | |
| Claude 4 | 2025-05 | Not disclosed | 200K | No | Strengthened coding and Agent scenarios | |
| Gemini (Google) | Gemini 1.5 Pro | 2024-05 | Not disclosed | 1M | No | Million-token native long context |
| Gemini 2.5 Pro | 2025-03 | Not disclosed | 1M | No | Reasoning + multimodal + ultra-long context | |
| Llama (Meta) | Llama 2 | 2023-07 | 7B / 13B / 70B | 4K | Yes | First large-scale commercially usable open model |
| Llama 3.1 | 2024-07 | 8B / 70B / 405B | 128K | Yes | 405B open-flagship + 128K context | |
| Llama 4 Scout | 2025-04 | 109B total / 17B active (MoE) | ~10M | Yes | Open MoE, ultra-long context | |
| Qwen (Alibaba) | Qwen2.5 | 2024-09 | 0.5B–72B | 128K | Yes | Full-size coverage, strong in Chinese + code |
| Qwen3 | 2025-04 | 0.6B–32B (dense) + MoE variant | ~256K | Yes | Hybrid reasoning (toggleable thinking mode) | |
| DeepSeek | DeepSeek-V3 | 2024-12 | 671B total / 37B active (MoE) | 128K | Yes (MIT) | Open MoE cost-performance benchmark |
| DeepSeek-R1 | 2025-01 | 671B + distilled small models | 128K | Yes (MIT) | First open reasoning model on par with closed models | |
| Mistral / Mixtral | Mistral 7B | 2023-09 | 7B | 32K | Yes (Apache 2.0) | A representative of small-model efficiency |
| Mixtral 8x7B | 2023-12 | 46.7B total / 12.9B active (MoE) | 32K | Yes (Apache 2.0) | The first breakout open-source MoE | |
| Mistral Large | 2024-02 | ~123B | 32K | No | European closed-flagship | |
| Gemma (Google) | Gemma 3 | 2025-03 | 1B / 4B / 12B / 27B | ~128K | Yes | Lightweight multimodal, edge-friendly |
| GLM (Zhipu) | GLM-4-9B | 2024-06 | 9B | 128K | Yes | Chinese conversation and tool calling |
| GLM-4.5 | 2025-02 | Not disclosed | ~128K | No | Flagship for Chinese language capability | |
| Phi (Microsoft) | Phi-4 | 2024-12 | 14B | ~16K | Yes (MIT) | "Textbook-quality data" for a small model |
| Command R (Cohere) | Command R+ | 2024-04 | 104B | 128K | Yes (CC-BY-NC, non-commercial) | Specially optimized for RAG scenarios |
| Grok (xAI) | Grok-1 | 2024-03 | 314B (MoE) | 8K | Yes | The largest open-weight model at its time |
| Kimi (Moonshot) | Kimi K2 | 2025-07 | 173B (MoE) | ~256K | Yes (Apache 2.0) | Ultra-long context + Agent scenarios |
| ERNIE (Baidu) | ERNIE 4.5 | 2025-06 | Not disclosed | ~128K | No | Chinese ecosystem and multimodal |
How to read this table
- "Parameters" and "context window" are the two columns that change most rapidly: after Llama 3.1 launched, the community widely adopted it as a baseline, and Qwen3/DeepSeek have since redefined open-source cost-performance. 2. Open-source licenses vary widely: MIT (Phi, DeepSeek, Mistral, some Qwen3 variants) is the most permissive; the Llama Community License and Gemma Terms come with additional conditions; Command R is non-commercial only. 3. For closed-source models with undisclosed parameters, capability can only be judged through evaluations and hands-on testing.
II. Categorization by Dimension
| Dimension | Category | Representative models | Notes |
|---|---|---|---|
| Open vs. Closed | Closed API | GPT-4o / GPT-5, Claude 4, Gemini 2.5 Pro, GLM-4.5 | Leading capability, zero ops overhead, but controlled, pay-per-use, and data-crossing needs evaluation |
| Open weights | Llama 3.1, Qwen3, DeepSeek-V3/R1, Mistral, Gemma, Phi, Kimi K2 | Controllable, privately deployable, capability catching up fast (gap narrowed rapidly post-2024) | |
| Architecture | Dense (fully connected) | Llama 3.1, Qwen2.5, Gemma, Phi | All parameters participate in computation; mainstream for small-to-mid sizes |
| MoE (sparse) | Mixtral, DeepSeek-V3, Llama 4, Qwen3-MoE, Kimi K2 | Large total parameters but only a subset activated per token; excellent cost-performance | |
| Modality | Text-only | Llama 3.1, DeepSeek-R1, Phi-4 | Pure text reasoning |
| Multimodal | GPT-4o, Gemini, Qwen2.5-VL, GLM-4V, Gemma 3 | Joint understanding of images/audio/video + text | |
| Reasoning mode | Standard / instant | GPT-4o, Claude, Qwen3 (thinking toggleable off) | Low latency, for everyday tasks |
| Reasoning-enhanced | o1/o3, DeepSeek-R1, Claude 3.7/4, Gemini 2.5 Pro | Long reasoning chains; stronger at math/code/complex reasoning but slower and more expensive |
III. Family at a Glance
A one-paragraph profile for each family, helping you build a mental map of "who's who and what they're good at."
OpenAI GPT series: GPT-1 (2018) introduced the "generative pretraining + fine-tuning" paradigm; GPT-3 (2020) pushed scale to 175B and demonstrated few-shot capability; InstructGPT/ChatGPT (2022) introduced RLHF alignment; GPT-4 (2023) achieved multimodal capability and exam-level performance; the o-series (2024) pioneered test-time reasoning scaling; GPT-5 (2025) natively fuses reasoning and tool use. Always an industry technology barometer. See The GPT Series.
Anthropic Claude: Founded by former OpenAI alignment researchers, known for "constitutional AI" and safety alignment. Claude 3 and beyond have consistently ranked in the top tier for long-context, coding, and Agent scenarios. Claude 3.7's hybrid reasoning (adjustable thinking budget) and Claude Code tooling have earned strong developer adoption.
Google Gemini: The native multimodal flagship after the merger of DeepMind and Google Brain. Gemini 1.5 was the first to push context windows to a million tokens; 2. Pro has caught up with the top tier in reasoning and multimodal understanding. The open-source Gemma family targets lightweight deployment from the same lineage.
Meta Llama: The de facto standard of the open-source ecosystem. Llama 2 (2023) was the first large-scale commercially usable open model. Llama 3.1 (2024) released a 405B open dense flagship with 128K context. Llama 4 shifts to MoE with 10-million-token context. A full ecosystem of fine-tuning, quantization, and deployment has formed around it. See Llama and the Open-Source Ecosystem.
Alibaba Qwen: The Chinese model family with the most comprehensive full-size open-source coverage (0.5B–72B). Qwen2.5 is one of the most popular bases for open fine-tuning and deployment; Qwen3 adds hybrid reasoning (toggleable thinking mode). Companion models Qwen2.5-Coder / Qwen2.5-VL cover code and multimodal.
DeepSeek: Known for "extreme engineering cost-performance": V2 validated MoE + MLA architecture, V3 (671B total / 37B active) achieved top-tier performance at a fraction of the training cost of comparable closed models, and R1 was the first open reasoning model to catch up to closed models. Training details are the most publicly documented, under the most permissive MIT license.
Mistral / Mixtral: The European open-source representative under the most permissive Apache 2.0 license. Mistral 7B is renowned for efficiency, Mixtral 8x7B was the first breakout open MoE, and Mistral Large is the closed flagship, with an advantage in European data-compliance (GDPR) scenarios. See MoE and Ultra-Large Models.
Google Gemma: Lightweight open models built on Gemini technology, ranging from 1B to 27B for edge to mid-range deployment. Gemma 3 supports multimodal and long context, suitable for consumer hardware and mobile devices.
Zhipu GLM: One of China's earliest open-source Chinese conversation models (ChatGLM-6B). The GLM-4 series is mature in Chinese conversation, tool calling, and Agent scenarios; GLM-4.5 is the flagship for Chinese language capability.
Microsoft Phi: Small model, big results — 1B–14B models trained on "textbook-quality" high-quality data. Phi-4 rivals models several times its size at coding and math. MIT-licensed, ideal for resource-constrained environments.
Cohere Command R: Optimized for enterprise RAG scenarios (retrieval citations, multilingual, tool calling). Command R+ at 104B is one of the few models specifically designed for "retrieval-augmented" use, though its license is non-commercial.
xAI Grok: Focuses on real-time data (from the X platform) and product integration. Grok-1 was once the largest open-weight model at 314B MoE; Grok-2/3 have shifted to closed-source and entered the top tier for reasoning.
Moonshot Kimi: Started with ultra-long context (one of the first to commercialize 200K context). Kimi K2 (2025.07) open-sourced a 173B MoE model targeting Agent scenarios.
Baidu ERNIE: The flagship behind ERNIE Bot. The Chinese ecosystem is comprehensive (search, cloud, agents). ERNIE 4.5 focuses on multimodal and Chinese understanding, available as a closed API.
IV. Ecosystem and Licensing Quick-Reference
The "license health check" before model selection — open weights ≠ free for commercial use:
| Model | License type | Commercial use? | Key notes |
|---|---|---|---|
| GPT / Claude / Gemini / GLM-4.5 / ERNIE | Closed API | Yes (pay-per-use) | Subject to terms of service; data-crossing requires evaluation |
| Llama series | Llama Community License | Yes (additional application if monthly active users exceed 700M) | Includes clauses prohibiting training competing models with outputs, etc. |
| Qwen2.5 / some Qwen3 variants | Apache 2.0 | Yes | Some derived versions may have additional terms |
| DeepSeek-V3 / R1 | MIT | Yes | One of the most permissive licenses |
| Mistral / Mixtral | Apache 2.0 | Yes | One of the most permissive licenses |
| Gemma series | Gemma License | Yes (with conditions) | Prohibits certain restricted uses — read the terms |
| GLM-4-9B and other open variants | Apache 2.0 | Yes | The closed flagship is not included |
| Phi series | MIT | Yes | One of the most permissive licenses |
| Command R+ | CC-BY-NC | No (non-commercial only) | Contact Cohere for commercial use |
| Grok-1 | Apache 2.0 | Yes | Grok-2/3 are closed-source |
| Kimi K2 | Apache 2.0 | Yes | Verify latest terms |
V. How to Choose a Model
There is no "best model," only the "right model for the scenario." Follow four steps: define the scenario → define constraints (budget/deployment/data compliance) → create a shortlist → test. The table below gives recommended starting points for common scenarios.
| Scenario | Recommendation | Reason |
|---|---|---|
| General conversation / writing (no deployment constraints) | GPT-4o / Claude 4 / Gemini 2.5 Flash | Strong overall capability, mature ecosystem of tools |
| Code generation and review | Claude 4 / GPT-5 (closed); Qwen2.5-Coder (open) | Long-leading on code benchmarks |
| Math and complex reasoning | o1 / Gemini 2.5 Pro (closed); DeepSeek-R1 (open) | Test-time scaling excels at multi-step derivation |
| Ultra-long documents (100K+ tokens) | Gemini 2.5 Pro / GPT-5 (closed); Kimi K2 / Llama 3.1 405B (open) | Native long context window |
| Chinese business deployment | Qwen3 / GLM-4.5 / DeepSeek | High Chinese corpus ratio, complete Chinese ecosystem |
| Private data + RAG | Command R+ / Qwen3 / Llama 3.1 | Retrieval-friendly or open-source, controllable, privately deployable |
| Local deployment (consumer GPU, 8–16 GB VRAM) | Qwen3-8B / Gemma 3 12B / Phi-4 | Quantized memory is manageable; see Deployment and Serving |
| Open fine-tuning base | Llama 3.1 8B / Qwen2.5-7B (small-to-mid); DeepSeek-V3 (large) | Most mature community ecosystem, tooling, and licensing |
| Cost-sensitive high-concurrency inference | DeepSeek-V3 (MoE) / quantized small models | Fewer active parameters, lower cost per token |
| Edge / on-device | Llama 3.2 1B/3B / Qwen3-0.6B / Gemma 3 1B | Millisecond latency, minimal VRAM |
1. Three Selection Case Studies
Case A: Mid-size company "internal knowledge base QA" (Chinese docs, private deployment, limited budget) → Choose Qwen3-32B (or Qwen2.5-14B) + vLLM + self-built RAG. Rationale: open-source and privately deployable, strong Chinese understanding, 14B/32B quantized VRAM is feasible (2–4 GPUs), and the Chinese community docs and toolchain are the most complete. Closed APIs can't work under "data must stay inside the corporate network" constraints — open weights are the only solution.
Case B: Code assistant product (high concurrency, low latency, API-acceptable) → Choose Claude 4 or GPT-5 API + prompt caching. Rationale: code capability is top-tier; the API eliminates GPU ops overhead. Route "latency-sensitive" small requests to a small model (e.g., GPT-4o mini or an open 8B), and reserve the large model for complex tasks only — model-tiered routing is the core cost-saving strategy.
Case C: Academic research team evaluating math/reasoning → Choose DeepSeek-R1 (open, reproducible, can run evaluations privately) or o3 API. Rationale: top-tier performance on reasoning benchmarks; the open version facilitates full experiment-environment logging, meeting academic reproducibility requirements.
Iron rules for model selection
- Get something working before optimizing: validate with an API first, then consider self-hosting. 2. Test with your own data: public leaderboards and your actual task may not share the same distribution. 3. Calculate total cost per token: compare closed API call volume against open-source self-hosting GPU depreciation — do the full math. 4. Refer to Framework and Tool Comparison for companion tooling and Deployment and Serving for resource estimation.
Three common misconceptions
- More parameters is always better: context length, task type, and latency budget all change the optimal answer; a 7B model with RAG often beats a 405B model trying to brute-force it.
- Open-source = free: GPUs, bandwidth, and ops personnel all cost money. Open-source just means a different cost structure.
- Go multimodal just because it's available: if you only use text, paying for multimodal capability is waste. Pick "good enough" rather than "most comprehensive."
VI. How to Access and Run
| Route | Suitable for | How |
|---|---|---|
| Closed API | No deployment infrastructure, seeking top capability, fast time-to-market | Register on each official platform and grab a key (OpenAI / Anthropic / Google / Zhipu / Moonshot, etc.), or use an aggregator platform like OpenRouter for a unified integration |
| Open weights | Private deployment, fine-tuning, cost optimization | Download from Hugging Face Hub → vLLM (serving) / llama.cpp (CPU edge) / Ollama (local experience) |
| Hybrid | Tiered routing | Simple tasks go to small models / open-source; complex tasks go to closed API |
Full memory estimation, quantization, concurrency metrics, and deployment steps are in Deployment and Serving. Tool selection is covered in Framework and Tool Comparison.
VII. Staying Current
The model market moves at a monthly pace — this compendium may already be lagging by the time you read it. We recommend pinning a few information sources and scanning them weekly:
- Official channels: OpenAI / Anthropic / Google / Meta's official blogs and model docs (see References).
- Community leaderboards: LMArena and Hugging Face trending models — check real-world model reputation.
- Papers: arXiv and frontier progress — see where the technology roadmap is heading next.
- Toolchain: vLLM / llama.cpp release notes — new model support typically appears here first.
Further Reading
- The GPT Series: From GPT-1 to GPT-4o — The full evolution of the GPT family
- Llama and the Open-Source Ecosystem — The competitive history of open-source models
- MoE and Ultra-Large Models — Profiles of Mixtral, DeepSeek-V3, and other MoE models
- ChatGPT and Conversational Models — Timelines for Claude / Gemini / ERNIE / Tongyi, and more
- Context and Long Context — The technology behind each model's context window
- Framework and Tool Comparison — How to build a toolchain once you've picked a model
- Deployment and Serving — Memory estimation, quantization, and inference optimization
References
- OpenAI model docs: https://platform.openai.com/docs/models
- Anthropic model docs: https://docs.anthropic.com/en/docs/about-claude/models/overview
- Google DeepMind Gemini: https://deepmind.google/models/gemini/
- Meta Llama official: https://www.llama.com/
- Qwen (Alibaba) GitHub: https://github.com/QwenLM/Qwen3
- DeepSeek-V3 GitHub: https://github.com/deepseek-ai/DeepSeek-V3
- DeepSeek-R1 GitHub: https://github.com/deepseek-ai/DeepSeek-R1
- Mistral AI official: https://mistral.ai/
- Google Gemma: https://ai.google.dev/gemma
- Zhipu GLM GitHub: https://github.com/THUDM/GLM-4
- Microsoft Phi-4 (HF): https://huggingface.co/microsoft/phi-4
- Kimi K2 GitHub: https://github.com/MoonshotAI/Kimi-K2
- Hugging Face model hub: https://huggingface.co/models