Skip to content

Model Compendium

At a glance quick-reference guide to mainstream large models: spec, licensing, and highlights comparison across GPT, Claude, Gemini, Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, Phi, and more — including a "how to choose a model" decision table and selection case studies.

This page contains time-sensitive content, current as of 2025-08; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

Model Compendium ​

This page is a quick-reference guide to mainstream large language models: family, release date, parameter count, context window, open-source status, and highlights — along with a categorization by dimension, a licensing quick-reference, and a "how to choose a model" decision table. Model specs and version rollouts move fast — this page's data is current as of August 2025; always refer to official sources for the latest information. Closed-source model parameters are generally undisclosed and marked as "not disclosed."

Model profiles become outdated quickly

GPT-4o was once the leader, and now there's the o-series and GPT-5. Llama went from version 2 to version 4 in just two years. This page serves only as an "introductory map" — always check the official docs before procurement or model selection (see References at the end). For evaluation scores, adopt the critical perspective outlined in Evaluation and Benchmarks rather than fixating on leaderboards.

I. Model Overview Table ​

Representative versions by family. "Open source" means weights are publicly available (note that licenses vary by family). Parameter counts are either officially published or based on industry consensus; "not disclosed" where the information is unavailable.

FamilyRepresentative versionRelease dateParametersContextOpen source?Highlights
GPT (OpenAI)GPT-32020-05175B2KNoFirst to demonstrate large-scale few-shot capability
GPT-42023-03Not disclosed (rumored MoE)8K / 32KNoExam-level capability, multimodal
GPT-4o2024-05Not disclosed128KNoNative multimodal, low-latency real-time interaction
o12024-09Not disclosed200KNoTest-time scaling (reasoning-time expansion)
GPT-52025-08Not disclosed~400KNoNative fusion of reasoning and tool use
Claude (Anthropic)Claude 3 series2024-03Not disclosed200KNoExcellent long-context and safety policy
Claude 3.7 Sonnet2025-02Not disclosed200KNoHybrid reasoning (controllable thinking budget)
Claude 42025-05Not disclosed200KNoStrengthened coding and Agent scenarios
Gemini (Google)Gemini 1.5 Pro2024-05Not disclosed1MNoMillion-token native long context
Gemini 2.5 Pro2025-03Not disclosed1MNoReasoning + multimodal + ultra-long context
Llama (Meta)Llama 22023-077B / 13B / 70B4KYesFirst large-scale commercially usable open model
Llama 3.12024-078B / 70B / 405B128KYes405B open-flagship + 128K context
Llama 4 Scout2025-04109B total / 17B active (MoE)~10MYesOpen MoE, ultra-long context
Qwen (Alibaba)Qwen2.52024-090.5B–72B128KYesFull-size coverage, strong in Chinese + code
Qwen32025-040.6B–32B (dense) + MoE variant~256KYesHybrid reasoning (toggleable thinking mode)
DeepSeekDeepSeek-V32024-12671B total / 37B active (MoE)128KYes (MIT)Open MoE cost-performance benchmark
DeepSeek-R12025-01671B + distilled small models128KYes (MIT)First open reasoning model on par with closed models
Mistral / MixtralMistral 7B2023-097B32KYes (Apache 2.0)A representative of small-model efficiency
Mixtral 8x7B2023-1246.7B total / 12.9B active (MoE)32KYes (Apache 2.0)The first breakout open-source MoE
Mistral Large2024-02~123B32KNoEuropean closed-flagship
Gemma (Google)Gemma 32025-031B / 4B / 12B / 27B~128KYesLightweight multimodal, edge-friendly
GLM (Zhipu)GLM-4-9B2024-069B128KYesChinese conversation and tool calling
GLM-4.52025-02Not disclosed~128KNoFlagship for Chinese language capability
Phi (Microsoft)Phi-42024-1214B~16KYes (MIT)"Textbook-quality data" for a small model
Command R (Cohere)Command R+2024-04104B128KYes (CC-BY-NC, non-commercial)Specially optimized for RAG scenarios
Grok (xAI)Grok-12024-03314B (MoE)8KYesThe largest open-weight model at its time
Kimi (Moonshot)Kimi K22025-07173B (MoE)~256KYes (Apache 2.0)Ultra-long context + Agent scenarios
ERNIE (Baidu)ERNIE 4.52025-06Not disclosed~128KNoChinese ecosystem and multimodal

How to read this table

  1. "Parameters" and "context window" are the two columns that change most rapidly: after Llama 3.1 launched, the community widely adopted it as a baseline, and Qwen3/DeepSeek have since redefined open-source cost-performance. 2. Open-source licenses vary widely: MIT (Phi, DeepSeek, Mistral, some Qwen3 variants) is the most permissive; the Llama Community License and Gemma Terms come with additional conditions; Command R is non-commercial only. 3. For closed-source models with undisclosed parameters, capability can only be judged through evaluations and hands-on testing.

II. Categorization by Dimension ​

DimensionCategoryRepresentative modelsNotes
Open vs. ClosedClosed APIGPT-4o / GPT-5, Claude 4, Gemini 2.5 Pro, GLM-4.5Leading capability, zero ops overhead, but controlled, pay-per-use, and data-crossing needs evaluation
Open weightsLlama 3.1, Qwen3, DeepSeek-V3/R1, Mistral, Gemma, Phi, Kimi K2Controllable, privately deployable, capability catching up fast (gap narrowed rapidly post-2024)
ArchitectureDense (fully connected)Llama 3.1, Qwen2.5, Gemma, PhiAll parameters participate in computation; mainstream for small-to-mid sizes
MoE (sparse)Mixtral, DeepSeek-V3, Llama 4, Qwen3-MoE, Kimi K2Large total parameters but only a subset activated per token; excellent cost-performance
ModalityText-onlyLlama 3.1, DeepSeek-R1, Phi-4Pure text reasoning
MultimodalGPT-4o, Gemini, Qwen2.5-VL, GLM-4V, Gemma 3Joint understanding of images/audio/video + text
Reasoning modeStandard / instantGPT-4o, Claude, Qwen3 (thinking toggleable off)Low latency, for everyday tasks
Reasoning-enhancedo1/o3, DeepSeek-R1, Claude 3.7/4, Gemini 2.5 ProLong reasoning chains; stronger at math/code/complex reasoning but slower and more expensive

III. Family at a Glance ​

A one-paragraph profile for each family, helping you build a mental map of "who's who and what they're good at."

OpenAI GPT series: GPT-1 (2018) introduced the "generative pretraining + fine-tuning" paradigm; GPT-3 (2020) pushed scale to 175B and demonstrated few-shot capability; InstructGPT/ChatGPT (2022) introduced RLHF alignment; GPT-4 (2023) achieved multimodal capability and exam-level performance; the o-series (2024) pioneered test-time reasoning scaling; GPT-5 (2025) natively fuses reasoning and tool use. Always an industry technology barometer. See The GPT Series.

Anthropic Claude: Founded by former OpenAI alignment researchers, known for "constitutional AI" and safety alignment. Claude 3 and beyond have consistently ranked in the top tier for long-context, coding, and Agent scenarios. Claude 3.7's hybrid reasoning (adjustable thinking budget) and Claude Code tooling have earned strong developer adoption.

Google Gemini: The native multimodal flagship after the merger of DeepMind and Google Brain. Gemini 1.5 was the first to push context windows to a million tokens; 2. Pro has caught up with the top tier in reasoning and multimodal understanding. The open-source Gemma family targets lightweight deployment from the same lineage.

Meta Llama: The de facto standard of the open-source ecosystem. Llama 2 (2023) was the first large-scale commercially usable open model. Llama 3.1 (2024) released a 405B open dense flagship with 128K context. Llama 4 shifts to MoE with 10-million-token context. A full ecosystem of fine-tuning, quantization, and deployment has formed around it. See Llama and the Open-Source Ecosystem.

Alibaba Qwen: The Chinese model family with the most comprehensive full-size open-source coverage (0.5B–72B). Qwen2.5 is one of the most popular bases for open fine-tuning and deployment; Qwen3 adds hybrid reasoning (toggleable thinking mode). Companion models Qwen2.5-Coder / Qwen2.5-VL cover code and multimodal.

DeepSeek: Known for "extreme engineering cost-performance": V2 validated MoE + MLA architecture, V3 (671B total / 37B active) achieved top-tier performance at a fraction of the training cost of comparable closed models, and R1 was the first open reasoning model to catch up to closed models. Training details are the most publicly documented, under the most permissive MIT license.

Mistral / Mixtral: The European open-source representative under the most permissive Apache 2.0 license. Mistral 7B is renowned for efficiency, Mixtral 8x7B was the first breakout open MoE, and Mistral Large is the closed flagship, with an advantage in European data-compliance (GDPR) scenarios. See MoE and Ultra-Large Models.

Google Gemma: Lightweight open models built on Gemini technology, ranging from 1B to 27B for edge to mid-range deployment. Gemma 3 supports multimodal and long context, suitable for consumer hardware and mobile devices.

Zhipu GLM: One of China's earliest open-source Chinese conversation models (ChatGLM-6B). The GLM-4 series is mature in Chinese conversation, tool calling, and Agent scenarios; GLM-4.5 is the flagship for Chinese language capability.

Microsoft Phi: Small model, big results — 1B–14B models trained on "textbook-quality" high-quality data. Phi-4 rivals models several times its size at coding and math. MIT-licensed, ideal for resource-constrained environments.

Cohere Command R: Optimized for enterprise RAG scenarios (retrieval citations, multilingual, tool calling). Command R+ at 104B is one of the few models specifically designed for "retrieval-augmented" use, though its license is non-commercial.

xAI Grok: Focuses on real-time data (from the X platform) and product integration. Grok-1 was once the largest open-weight model at 314B MoE; Grok-2/3 have shifted to closed-source and entered the top tier for reasoning.

Moonshot Kimi: Started with ultra-long context (one of the first to commercialize 200K context). Kimi K2 (2025.07) open-sourced a 173B MoE model targeting Agent scenarios.

Baidu ERNIE: The flagship behind ERNIE Bot. The Chinese ecosystem is comprehensive (search, cloud, agents). ERNIE 4.5 focuses on multimodal and Chinese understanding, available as a closed API.

IV. Ecosystem and Licensing Quick-Reference ​

The "license health check" before model selection — open weights ≠ free for commercial use:

ModelLicense typeCommercial use?Key notes
GPT / Claude / Gemini / GLM-4.5 / ERNIEClosed APIYes (pay-per-use)Subject to terms of service; data-crossing requires evaluation
Llama seriesLlama Community LicenseYes (additional application if monthly active users exceed 700M)Includes clauses prohibiting training competing models with outputs, etc.
Qwen2.5 / some Qwen3 variantsApache 2.0YesSome derived versions may have additional terms
DeepSeek-V3 / R1MITYesOne of the most permissive licenses
Mistral / MixtralApache 2.0YesOne of the most permissive licenses
Gemma seriesGemma LicenseYes (with conditions)Prohibits certain restricted uses — read the terms
GLM-4-9B and other open variantsApache 2.0YesThe closed flagship is not included
Phi seriesMITYesOne of the most permissive licenses
Command R+CC-BY-NCNo (non-commercial only)Contact Cohere for commercial use
Grok-1Apache 2.0YesGrok-2/3 are closed-source
Kimi K2Apache 2.0YesVerify latest terms

V. How to Choose a Model ​

There is no "best model," only the "right model for the scenario." Follow four steps: define the scenario → define constraints (budget/deployment/data compliance) → create a shortlist → test. The table below gives recommended starting points for common scenarios.

ScenarioRecommendationReason
General conversation / writing (no deployment constraints)GPT-4o / Claude 4 / Gemini 2.5 FlashStrong overall capability, mature ecosystem of tools
Code generation and reviewClaude 4 / GPT-5 (closed); Qwen2.5-Coder (open)Long-leading on code benchmarks
Math and complex reasoningo1 / Gemini 2.5 Pro (closed); DeepSeek-R1 (open)Test-time scaling excels at multi-step derivation
Ultra-long documents (100K+ tokens)Gemini 2.5 Pro / GPT-5 (closed); Kimi K2 / Llama 3.1 405B (open)Native long context window
Chinese business deploymentQwen3 / GLM-4.5 / DeepSeekHigh Chinese corpus ratio, complete Chinese ecosystem
Private data + RAGCommand R+ / Qwen3 / Llama 3.1Retrieval-friendly or open-source, controllable, privately deployable
Local deployment (consumer GPU, 8–16 GB VRAM)Qwen3-8B / Gemma 3 12B / Phi-4Quantized memory is manageable; see Deployment and Serving
Open fine-tuning baseLlama 3.1 8B / Qwen2.5-7B (small-to-mid); DeepSeek-V3 (large)Most mature community ecosystem, tooling, and licensing
Cost-sensitive high-concurrency inferenceDeepSeek-V3 (MoE) / quantized small modelsFewer active parameters, lower cost per token
Edge / on-deviceLlama 3.2 1B/3B / Qwen3-0.6B / Gemma 3 1BMillisecond latency, minimal VRAM

1. Three Selection Case Studies ​

Case A: Mid-size company "internal knowledge base QA" (Chinese docs, private deployment, limited budget) → Choose Qwen3-32B (or Qwen2.5-14B) + vLLM + self-built RAG. Rationale: open-source and privately deployable, strong Chinese understanding, 14B/32B quantized VRAM is feasible (2–4 GPUs), and the Chinese community docs and toolchain are the most complete. Closed APIs can't work under "data must stay inside the corporate network" constraints — open weights are the only solution.

Case B: Code assistant product (high concurrency, low latency, API-acceptable) → Choose Claude 4 or GPT-5 API + prompt caching. Rationale: code capability is top-tier; the API eliminates GPU ops overhead. Route "latency-sensitive" small requests to a small model (e.g., GPT-4o mini or an open 8B), and reserve the large model for complex tasks only — model-tiered routing is the core cost-saving strategy.

Case C: Academic research team evaluating math/reasoning → Choose DeepSeek-R1 (open, reproducible, can run evaluations privately) or o3 API. Rationale: top-tier performance on reasoning benchmarks; the open version facilitates full experiment-environment logging, meeting academic reproducibility requirements.

Iron rules for model selection

  1. Get something working before optimizing: validate with an API first, then consider self-hosting. 2. Test with your own data: public leaderboards and your actual task may not share the same distribution. 3. Calculate total cost per token: compare closed API call volume against open-source self-hosting GPU depreciation — do the full math. 4. Refer to Framework and Tool Comparison for companion tooling and Deployment and Serving for resource estimation.

Three common misconceptions

  • More parameters is always better: context length, task type, and latency budget all change the optimal answer; a 7B model with RAG often beats a 405B model trying to brute-force it.
  • Open-source = free: GPUs, bandwidth, and ops personnel all cost money. Open-source just means a different cost structure.
  • Go multimodal just because it's available: if you only use text, paying for multimodal capability is waste. Pick "good enough" rather than "most comprehensive."

VI. How to Access and Run ​

RouteSuitable forHow
Closed APINo deployment infrastructure, seeking top capability, fast time-to-marketRegister on each official platform and grab a key (OpenAI / Anthropic / Google / Zhipu / Moonshot, etc.), or use an aggregator platform like OpenRouter for a unified integration
Open weightsPrivate deployment, fine-tuning, cost optimizationDownload from Hugging Face Hub → vLLM (serving) / llama.cpp (CPU edge) / Ollama (local experience)
HybridTiered routingSimple tasks go to small models / open-source; complex tasks go to closed API

Full memory estimation, quantization, concurrency metrics, and deployment steps are in Deployment and Serving. Tool selection is covered in Framework and Tool Comparison.

VII. Staying Current ​

The model market moves at a monthly pace — this compendium may already be lagging by the time you read it. We recommend pinning a few information sources and scanning them weekly:

  1. Official channels: OpenAI / Anthropic / Google / Meta's official blogs and model docs (see References).
  2. Community leaderboards: LMArena and Hugging Face trending models — check real-world model reputation.
  3. Papers: arXiv and frontier progress — see where the technology roadmap is heading next.
  4. Toolchain: vLLM / llama.cpp release notes — new model support typically appears here first.

Further Reading ​

References ​