Appearance
Models and Leaderboards Cheat Sheet
This page is the living "model selection" sheet of the AI Hot Concepts Handbook: one table to see the major closed- and open-source models at a glance, how to read leaderboards without falling into common traps, and an answer to "which model should I actually use for my scenario?" Data is current as of mid-2025 (
dataAsOf: 2025-06in the frontmatter).
Models iterate far faster than any handbook can be updated; every parameter, price, and leaderboard rank on this page is a "snapshot" — always verify against official sources before relying on it. In the LLM world, three months make a generation: GPT spawned the o-series, Claude adopted hybrid reasoning, Qwen jumped from 2.5 to 3 within a single year, and the top of the leaderboards changes hands almost monthly. So this page is not meant to be an "authoritative database" — its goal is to help you build a selection framework: which dimensions to look at, which kinds of leaderboards to trust, and what logic to apply when choosing.
Timeliness Disclaimer
Information on this page is current as of 2025-06. Model versions (GPT-5, subsequent Llama 4 releases, etc.), prices, context windows, and leaderboard rankings may all have changed since. Before making any technical decision, defer to the official documentation. The sections covering general principles don't go stale — read those with confidence.
Closed-Source Model Comparison
Closed-source models (exposed via APIs) currently represent the "capability ceiling". Compare them along five dimensions: capability positioning, context window, access model, price range, ecosystem and tooling.
| Model | Vendor | Capability Profile | Context Window | Access | Price Range (per 1M tokens) |
|---|---|---|---|---|---|
| GPT-4o | OpenAI | All-round multimodal (text/image/audio) with mature instruction following | 128K | API / web / app | ~$2.5 input, ~$10 output |
| o3 / o4-mini | OpenAI | Reasoning models: long chains of thought for math, code, and complex planning | ~200K (o3) / ~200K (o4-mini) | API | Above GPT-4o; starts roughly in the $2–$11 range |
| Claude 3.5/3.7 Sonnet | Anthropic | 3.5 is all-round and dependable; 3.7 pioneered "hybrid reasoning" (instant answers + extended thinking), strong at code and agents | 200K | API / Claude apps | ~$3 input, ~$15 output |
| Gemini 2.x / 2.5 Pro | Natively multimodal + ultra-long context (2.5 Pro is 1M, 2M available on request), enhanced reasoning | 1M (2.5 Pro) | API / Gemini apps | ~$1.25 input, ~$10 output | |
| DeepSeek (V3/R1 API) | DeepSeek | Strong Chinese, exceptional value for money; R1 is a reasoning model | ~128K | Dual channel: API + open weights | Well below mainstream Western models (R1 ~$0.55/$2.19) |
| Doubao 1.5/2.0 | ByteDance | Chinese conversation, multimodal, tool calling; strong ToB ecosystem | Long-text support | API (Volcano Engine) | RMB-based pay-as-you-go pricing, competitive |
| Qwen-Max/Turbo | Alibaba Cloud | Strong Chinese, balanced multimodal and agent capabilities, OpenAI-compatible interface | Long-text support | API (Bailian platform) | RMB pricing, often with free quota |
| ERNIE 4.x | Baidu | Chinese understanding and generation, knowledge-enhanced, tied to the Baidu ecosystem | Long-text support | API / web; ERNIE 4.5 was open-sourced in 2025-06 | RMB pricing |
Quick Take
Don't start by comparing specs — compare what you need: for the strongest reasoning, pick o3/Claude 3.7; for ultra-long-context document analysis, pick Gemini 2.5 Pro; for affordable, high-volume Chinese, pick DeepSeek or one of the three Chinese providers. For the general comparison logic behind closed-source models, see Large Language Models.
Open-Source Model Comparison
Open-source models (open weights) stand for "controllability": local deployment, fine-tuning, and no API vendor lock-in. Evaluate them on parameter count, license, distinctive strengths, ecosystem.
| Model | Parameters | License | Distinctive Strengths | Ecosystem |
|---|---|---|---|---|
| Llama 3.1/3.2/3.3/4 | 3.3 centers on 70B (3.1 flagship is 405B); Llama 4 Scout ~109B (17B active) / Maverick ~400B | Llama license (commercial use allowed; special authorization required above 700M monthly active users) | The broadest ecosystem worldwide: the most complete fine-tuning toolchains, quantization, and deployment docs | A massive number of derivatives — almost "everything can be a Llama" |
| Qwen 2.5 / Qwen3 | 2.5 from 0.5B to 72B; Qwen3 from 0.6B to 235B (including MoE) | Apache-2.0 | The strongest open-source Chinese capability; Qwen3 introduces a hybrid reasoning mode (fast/slow thinking) | Available on both ModelScope and Hugging Face, with an active Chinese community |
| DeepSeek-V3 / R1 | Both V3 and R1 are 671B (MoE, ~37B active) | MIT | Cost-effective MoE architecture; R1 is the benchmark for open-source reasoning models and drew worldwide attention | The R1-Distill family of distilled small models suits local deployment |
| Mistral / Mixtral | 7B, 8x7B, 8x22B, plus Small (23B) / Medium (123B) | Apache-2.0 | Europe's open-source flagship and an MoE pioneer; Codestral for code | Well supported by European clouds and self-hosting platforms |
| Phi-3 / Phi-4 | 3.8B / 7B / 14B (Phi-3); Phi-4 is 14B | MIT | Strong reasoning at small parameter counts, trained on "textbook-quality" data, low resource footprint | Friendly to edge devices, teaching demos, and local deployment |
Quick Take
For Chinese, pick Qwen; to ride the widest ecosystem, pick Llama; for value and reasoning, pick DeepSeek; for small models, pick Phi. The biggest value of open-source models is fine-tunability — use peft + trl to bake your business data into the model, something closed-source APIs can't offer.
How to Read the Leaderboards
Three Main Types of Leaderboards
| Leaderboard | Mechanism | What to Look At | What to Watch Out For |
|---|---|---|---|
| LMSYS Chatbot Arena | Blind user voting + Elo scoring | An "opinion poll" of overall conversational experience | Reflects average impressions, not specific tasks; the voter mix skews results |
| LLM Leaderboard (formerly Open LLM Leaderboard) | Automated runs of open-source models on standardized benchmarks | Reproducible side-by-side comparison of open-source models | Now upgraded to 2.0, which switched to instruction-following evals (IFEval, BBH, MATH, etc.); old-board scores are not directly comparable |
| MMLU / GPQA / HumanEval and other benchmark boards | Automatic scoring on fixed question banks | Knowledge breadth (MMLU), research-grade reasoning (GPQA), code (HumanEval) | Single benchmarks suffer from saturation and "benchmark gaming"; see LLM Evaluation and Benchmarks |
The Landscape You'd Most Likely See as of Mid-2025
- Arena leaders: GPT-4o/o3, Claude 3.7 Sonnet, and Gemini 2.5 Pro take turns at the top of the Elo board, with gaps within the margin of error — "who's #1" comes down to your style preferences.
- On the open-source side: Qwen3 and DeepSeek-R1 are the "two Chinese champions", frequently going toe-to-toe with top closed-source models on Chinese and reasoning leaderboards; Llama 4's reception after launch was mixed, but its ecosystem remains its biggest moat.
- Reasoning models became the new mainline: o3, DeepSeek-R1, Claude 3.7 (thinking mode), and Qwen3 (hybrid reasoning) all treat "thinking longer" as a capability growth lever; math and code leaderboards are now almost entirely swept by reasoning models.
The "Leaderboard ≠ Suitability" Warning
Never Choose a Model by Leaderboard Rank Alone
Ranking first doesn't mean it's right for you. Leaderboards only measure "average question-answering ability" and can't measure three things: (1) performance on your domain data (legal/medical/your private documents); (2) cost and latency (the priciest, fastest model handling customer support tickets is waste); (3) privacy and compliance (can the data leave your domain?). The right approach: use leaderboards to shortlist candidates, then let your real business evaluation set make the final call — see Building an LLM Evaluation Suite.
Model Selection by Scenario
| Scenario | First Choice | Alternatives | Notes |
|---|---|---|---|
| Chinese conversation / customer support | Qwen3, DeepSeek | Doubao, Qwen | Wins on both Chinese corpus quality and price |
| Code generation and review | Claude 3.7 Sonnet | GPT-4o / o4-mini, Qwen2.5-Coder | Long context + code understanding; see the Copilot case study |
| Long-document RAG | Gemini 2.5 Pro (1M context) | Claude 3.7 (200K), Qwen long-context editions | Ultra-long context cuts chunking overhead; see Build a RAG App from Scratch |
| Agent tool calling | GPT-4o / Claude 3.7 | Qwen3, DeepSeek | Tool-calling reliability comes first; see AI Agents and Build an Agent from Scratch |
| Edge / local deployment | Phi-4, Qwen2.5-3B | Llama 3.2 1B/3B | Parameter count and VRAM are hard constraints; Deployment and Inference Optimization in Practice is a complete guide |
Embedding and Vision Models at a Glance
Models aren't just "generative". RAG retrieval quality and multimodal image-text alignment rely on a different class of specialized models.
| Type | Representative | Highlights | License / Access | Related Concept |
|---|---|---|---|---|
| Embedding | BAAI bge-m3 | Strong Chinese/multilingual retrieval; the go-to open-source embedding for RAG | Open source (MIT) | Vector Databases and Semantic Search |
| Embedding | OpenAI text-embedding-3 | English and multilingual; a stable, worry-free API | Closed API | Vector Databases and Semantic Search |
| Embedding | Cohere embed-v3 | Multilingual; pairs with rerank models | Closed API | Vector Databases and Semantic Search |
| Vision-language alignment | CLIP (OpenAI) | The foundational model trained on image-text pairs; both retrieval and generation build on it | Open source | Multimodal Models |
| Vision-language alignment | SigLIP (Google) | An improved CLIP with a more stable contrastive loss | Open source | Multimodal Models |
Quick Take
Pick embedding models by language and domain (bge first for Chinese corpora; any API works for English), and pick vision models based on whether you want "image-text retrieval" or "image understanding" — the former is CLIP/SigLIP territory; for the latter, go straight to a multimodal LLM.
Where to Get Models
| Channel | Best For | Notes |
|---|---|---|
| Hugging Face | Developers worldwide | One-stop hosting for open weights and datasets; one line of transformers loads a model; mirror sites help if you're in mainland China |
| ModelScope | Developers in mainland China | Built by Alibaba; the fullest collection of Chinese models (where Qwen officially debuts) with fast domestic downloads |
| Official API consoles | Application developers | OpenAI / Anthropic / Google / Volcano Engine / Alibaba Bailian / Baidu Qianfan — pay-as-you-go with zero ops |
| Ollama library | Local tinkerers | Run open-source models with a single command; ideal for personal experimentation and prototypes |
| vLLM + self-hosting | Production teams | Open weights + high-throughput inference, data never leaves your perimeter; see Inference Optimization and Quantization |
Self-Hosting Is Not Free
Open-source model weights cost nothing, but GPUs, electricity, and ops all cost money. A 70B+ model won't run on a single machine, and multi-GPU setups get expensive — factor "total cost of ownership" into your selection; it's often the real decision point between open source and closed source.
Further Reading
- Datasets and Tools Directory — which data to use for training and evaluation once you've picked a model
- LLM Evaluation and Benchmarks — the methodology and common pitfalls behind the leaderboards
- Large Language Models (LLM) — the capability map and underlying principles
- Fine-Tuning and PEFT (LoRA) — how open-source models become your specialized model
- AI Agents — how to choose on tool-calling capability
- Deployment and Inference Optimization in Practice — local deployment and cost control
- Curated Resource List and Glossary — more resources and terminology
- Learning Paths — fitting model selection into your overall learning route
References
- OpenAI API docs — official specs and pricing for GPT-4o / the o-series
- Anthropic model docs — Claude 3.5/3.7 capabilities and pricing
- Google Gemini models page — Gemini 2.x context windows and pricing
- DeepSeek API pricing — official pricing and context for V3 / R1
- Qwen official GitHub — release notes for the open-source Qwen series
- LMSYS Chatbot Arena — live data for the Elo voting leaderboard
- Hugging Face LLM Leaderboard — open-source model evaluation leaderboard
- MMLU paper and GPQA paper — original sources of the benchmarks
- bge model page — official description of the open-source embedding model
- CLIP paper — the original image-text alignment paper (Radford et al., 2021)
Model parameters, prices, and rankings on this page were compiled from each model's official channels as of mid-2025; always double-check the latest official documentation before use.