Skip to content

Models and Leaderboards Cheat Sheet

At a glance A quick-reference guide to models and leaderboards — comparison tables of the major closed- and open-source LLMs as of mid-2025, how to read leaderboard rankings without being misled (rank is not suitability), scenario-based selection guidance, plus embedding models, vision models, and where to get them.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Models and Leaderboards Cheat Sheet ​

This page is the living "model selection" sheet of the AI Hot Concepts Handbook: one table to see the major closed- and open-source models at a glance, how to read leaderboards without falling into common traps, and an answer to "which model should I actually use for my scenario?" Data is current as of mid-2025 (dataAsOf: 2025-06 in the frontmatter).

Models iterate far faster than any handbook can be updated; every parameter, price, and leaderboard rank on this page is a "snapshot" — always verify against official sources before relying on it. In the LLM world, three months make a generation: GPT spawned the o-series, Claude adopted hybrid reasoning, Qwen jumped from 2.5 to 3 within a single year, and the top of the leaderboards changes hands almost monthly. So this page is not meant to be an "authoritative database" — its goal is to help you build a selection framework: which dimensions to look at, which kinds of leaderboards to trust, and what logic to apply when choosing.

Timeliness Disclaimer

Information on this page is current as of 2025-06. Model versions (GPT-5, subsequent Llama 4 releases, etc.), prices, context windows, and leaderboard rankings may all have changed since. Before making any technical decision, defer to the official documentation. The sections covering general principles don't go stale — read those with confidence.

Closed-Source Model Comparison ​

Closed-source models (exposed via APIs) currently represent the "capability ceiling". Compare them along five dimensions: capability positioning, context window, access model, price range, ecosystem and tooling.

ModelVendorCapability ProfileContext WindowAccessPrice Range (per 1M tokens)
GPT-4oOpenAIAll-round multimodal (text/image/audio) with mature instruction following128KAPI / web / app~$2.5 input, ~$10 output
o3 / o4-miniOpenAIReasoning models: long chains of thought for math, code, and complex planning~200K (o3) / ~200K (o4-mini)APIAbove GPT-4o; starts roughly in the $2–$11 range
Claude 3.5/3.7 SonnetAnthropic3.5 is all-round and dependable; 3.7 pioneered "hybrid reasoning" (instant answers + extended thinking), strong at code and agents200KAPI / Claude apps~$3 input, ~$15 output
Gemini 2.x / 2.5 ProGoogleNatively multimodal + ultra-long context (2.5 Pro is 1M, 2M available on request), enhanced reasoning1M (2.5 Pro)API / Gemini apps~$1.25 input, ~$10 output
DeepSeek (V3/R1 API)DeepSeekStrong Chinese, exceptional value for money; R1 is a reasoning model~128KDual channel: API + open weightsWell below mainstream Western models (R1 ~$0.55/$2.19)
Doubao 1.5/2.0ByteDanceChinese conversation, multimodal, tool calling; strong ToB ecosystemLong-text supportAPI (Volcano Engine)RMB-based pay-as-you-go pricing, competitive
Qwen-Max/TurboAlibaba CloudStrong Chinese, balanced multimodal and agent capabilities, OpenAI-compatible interfaceLong-text supportAPI (Bailian platform)RMB pricing, often with free quota
ERNIE 4.xBaiduChinese understanding and generation, knowledge-enhanced, tied to the Baidu ecosystemLong-text supportAPI / web; ERNIE 4.5 was open-sourced in 2025-06RMB pricing

Quick Take

Don't start by comparing specs — compare what you need: for the strongest reasoning, pick o3/Claude 3.7; for ultra-long-context document analysis, pick Gemini 2.5 Pro; for affordable, high-volume Chinese, pick DeepSeek or one of the three Chinese providers. For the general comparison logic behind closed-source models, see Large Language Models.

Open-Source Model Comparison ​

Open-source models (open weights) stand for "controllability": local deployment, fine-tuning, and no API vendor lock-in. Evaluate them on parameter count, license, distinctive strengths, ecosystem.

ModelParametersLicenseDistinctive StrengthsEcosystem
Llama 3.1/3.2/3.3/43.3 centers on 70B (3.1 flagship is 405B); Llama 4 Scout ~109B (17B active) / Maverick ~400BLlama license (commercial use allowed; special authorization required above 700M monthly active users)The broadest ecosystem worldwide: the most complete fine-tuning toolchains, quantization, and deployment docsA massive number of derivatives — almost "everything can be a Llama"
Qwen 2.5 / Qwen32.5 from 0.5B to 72B; Qwen3 from 0.6B to 235B (including MoE)Apache-2.0The strongest open-source Chinese capability; Qwen3 introduces a hybrid reasoning mode (fast/slow thinking)Available on both ModelScope and Hugging Face, with an active Chinese community
DeepSeek-V3 / R1Both V3 and R1 are 671B (MoE, ~37B active)MITCost-effective MoE architecture; R1 is the benchmark for open-source reasoning models and drew worldwide attentionThe R1-Distill family of distilled small models suits local deployment
Mistral / Mixtral7B, 8x7B, 8x22B, plus Small (23B) / Medium (123B)Apache-2.0Europe's open-source flagship and an MoE pioneer; Codestral for codeWell supported by European clouds and self-hosting platforms
Phi-3 / Phi-43.8B / 7B / 14B (Phi-3); Phi-4 is 14BMITStrong reasoning at small parameter counts, trained on "textbook-quality" data, low resource footprintFriendly to edge devices, teaching demos, and local deployment

Quick Take

For Chinese, pick Qwen; to ride the widest ecosystem, pick Llama; for value and reasoning, pick DeepSeek; for small models, pick Phi. The biggest value of open-source models is fine-tunability — use peft + trl to bake your business data into the model, something closed-source APIs can't offer.

How to Read the Leaderboards ​

Three Main Types of Leaderboards ​

LeaderboardMechanismWhat to Look AtWhat to Watch Out For
LMSYS Chatbot ArenaBlind user voting + Elo scoringAn "opinion poll" of overall conversational experienceReflects average impressions, not specific tasks; the voter mix skews results
LLM Leaderboard (formerly Open LLM Leaderboard)Automated runs of open-source models on standardized benchmarksReproducible side-by-side comparison of open-source modelsNow upgraded to 2.0, which switched to instruction-following evals (IFEval, BBH, MATH, etc.); old-board scores are not directly comparable
MMLU / GPQA / HumanEval and other benchmark boardsAutomatic scoring on fixed question banksKnowledge breadth (MMLU), research-grade reasoning (GPQA), code (HumanEval)Single benchmarks suffer from saturation and "benchmark gaming"; see LLM Evaluation and Benchmarks

The Landscape You'd Most Likely See as of Mid-2025 ​

  • Arena leaders: GPT-4o/o3, Claude 3.7 Sonnet, and Gemini 2.5 Pro take turns at the top of the Elo board, with gaps within the margin of error — "who's #1" comes down to your style preferences.
  • On the open-source side: Qwen3 and DeepSeek-R1 are the "two Chinese champions", frequently going toe-to-toe with top closed-source models on Chinese and reasoning leaderboards; Llama 4's reception after launch was mixed, but its ecosystem remains its biggest moat.
  • Reasoning models became the new mainline: o3, DeepSeek-R1, Claude 3.7 (thinking mode), and Qwen3 (hybrid reasoning) all treat "thinking longer" as a capability growth lever; math and code leaderboards are now almost entirely swept by reasoning models.

The "Leaderboard ≠ Suitability" Warning ​

Never Choose a Model by Leaderboard Rank Alone

Ranking first doesn't mean it's right for you. Leaderboards only measure "average question-answering ability" and can't measure three things: (1) performance on your domain data (legal/medical/your private documents); (2) cost and latency (the priciest, fastest model handling customer support tickets is waste); (3) privacy and compliance (can the data leave your domain?). The right approach: use leaderboards to shortlist candidates, then let your real business evaluation set make the final call — see Building an LLM Evaluation Suite.

Model Selection by Scenario ​

ScenarioFirst ChoiceAlternativesNotes
Chinese conversation / customer supportQwen3, DeepSeekDoubao, QwenWins on both Chinese corpus quality and price
Code generation and reviewClaude 3.7 SonnetGPT-4o / o4-mini, Qwen2.5-CoderLong context + code understanding; see the Copilot case study
Long-document RAGGemini 2.5 Pro (1M context)Claude 3.7 (200K), Qwen long-context editionsUltra-long context cuts chunking overhead; see Build a RAG App from Scratch
Agent tool callingGPT-4o / Claude 3.7Qwen3, DeepSeekTool-calling reliability comes first; see AI Agents and Build an Agent from Scratch
Edge / local deploymentPhi-4, Qwen2.5-3BLlama 3.2 1B/3BParameter count and VRAM are hard constraints; Deployment and Inference Optimization in Practice is a complete guide

Embedding and Vision Models at a Glance ​

Models aren't just "generative". RAG retrieval quality and multimodal image-text alignment rely on a different class of specialized models.

TypeRepresentativeHighlightsLicense / AccessRelated Concept
EmbeddingBAAI bge-m3Strong Chinese/multilingual retrieval; the go-to open-source embedding for RAGOpen source (MIT)Vector Databases and Semantic Search
EmbeddingOpenAI text-embedding-3English and multilingual; a stable, worry-free APIClosed APIVector Databases and Semantic Search
EmbeddingCohere embed-v3Multilingual; pairs with rerank modelsClosed APIVector Databases and Semantic Search
Vision-language alignmentCLIP (OpenAI)The foundational model trained on image-text pairs; both retrieval and generation build on itOpen sourceMultimodal Models
Vision-language alignmentSigLIP (Google)An improved CLIP with a more stable contrastive lossOpen sourceMultimodal Models

Quick Take

Pick embedding models by language and domain (bge first for Chinese corpora; any API works for English), and pick vision models based on whether you want "image-text retrieval" or "image understanding" — the former is CLIP/SigLIP territory; for the latter, go straight to a multimodal LLM.

Where to Get Models ​

ChannelBest ForNotes
Hugging FaceDevelopers worldwideOne-stop hosting for open weights and datasets; one line of transformers loads a model; mirror sites help if you're in mainland China
ModelScopeDevelopers in mainland ChinaBuilt by Alibaba; the fullest collection of Chinese models (where Qwen officially debuts) with fast domestic downloads
Official API consolesApplication developersOpenAI / Anthropic / Google / Volcano Engine / Alibaba Bailian / Baidu Qianfan — pay-as-you-go with zero ops
Ollama libraryLocal tinkerersRun open-source models with a single command; ideal for personal experimentation and prototypes
vLLM + self-hostingProduction teamsOpen weights + high-throughput inference, data never leaves your perimeter; see Inference Optimization and Quantization

Self-Hosting Is Not Free

Open-source model weights cost nothing, but GPUs, electricity, and ops all cost money. A 70B+ model won't run on a single machine, and multi-GPU setups get expensive — factor "total cost of ownership" into your selection; it's often the real decision point between open source and closed source.

Further Reading ​

References ​

Model parameters, prices, and rankings on this page were compiled from each model's official channels as of mid-2025; always double-check the latest official documentation before use.