Appearance
DeepSeek-R1 and Reasoning Models
A reasoning model is a large language model (LLM) that produces an internal chain-of-thought (CoT) before generating its final answer — think it through first, then respond. Rather than blurting out a conclusion the way an ordinary LLM does, it writes its "thinking process" explicitly into the generation sequence. In January 2025, DeepSeek released and open-sourced the reasoning model DeepSeek-R1, which — at a training cost reportedly far below that of comparable models — matched or even beat OpenAI's o1 on reasoning tasks such as math and code. The release turned the "reasoning model" from a laboratory term into a global talking point and reshaped the fundamentals of open versus closed models, cost versus capability, training paradigms, and agent deployment.
Using R1 as our lens, this article makes three things clear: what actually separates reasoning models from ordinary LLMs, which exceptionally cost-effective techniques R1 employed, and the ripple effects it triggered across the industry. If you want to brush up on the basics first, read alongside Large Language Models (LLMs) and Prompt Engineering.
1. Background: A Global "Low-Cost, High-Performance" Phenomenon
1.1 Timeline and How It Unfolded
| Date | Event | Significance |
|---|---|---|
| 2024-09 | OpenAI releases o1-preview / o1-mini | The "reasoning model" category is born, showcasing chain-of-thought reasoning |
| 2024-12 | OpenAI releases o3 (with an o3-mini preview) | Reasoning takes another step up, but remains closed-source |
| 2025-01-20 | DeepSeek-R1 officially released and open-sourced (MIT-licensed weights) | o1-level reasoning, openly available |
| 2025-02 | DeepSeek-R1-V (multimodal reasoning) and a 167B MoE reasoning variant released | Reasoning extends to multimodality and broader deployment |
The day R1 launched, it ignited a global media storm: downloads of its weights on Hugging Face surged within hours, multiple cloud providers (Nvidia, Microsoft, Amazon, and others) announced support, and the app briefly topped the App Store free charts. For an open-source model from a Chinese startup, that level of attention was unprecedented — the industry widely dubbed it "AI's Sputnik moment."
1.2 The DeepSeek Story and the "Low-Cost, High-Performance" Narrative
What made R1 stunning was not that it was the most capable model, but its price-to-performance ratio. According to the technical report and industry estimates, R1's training cost was far below what OpenAI's comparable models are rumored to have cost:
- RL-stage compute: the report states that R1-Zero's reinforcement learning training used only about 8,000 V100 GPU hours — typically the budget for a small experiment at a big tech company;
- Overall training cost estimate: taking DeepSeek-V3's reported ~2.78 million H800 GPU hours and roughly US$5.57 million (industry estimate) as a reference, R1's full training cost is widely believed to be in the low millions of dollars — versus rumored training budgets of hundreds of millions of dollars for closed models, a gap of two to three orders of magnitude;
- Inference and deployment costs: R1 uses a Mixture-of-Experts (MoE) architecture with 671B total parameters but only about 37B activated per inference, and with the quantization and inference optimizations the open-source community adapted quickly (see Inference Optimization and Quantization), the cost per query was driven extremely low.
"Algorithmic innovation can partially make up for a compute gap" — before R1 this was a hypothesis; after R1 it became industry consensus. It also explains why the DeepSeek family has been a permanent fixture on the Model and Leaderboard Quick Reference page ever since.
2. What Is a Reasoning Model
2.1 The Mechanism: Making "Thinking" Part of Generation
An ordinary LLM works by input → answer, directly: ask it the classic puzzle "Chickens and rabbits are shut in a cage; together they have 35 heads and 94 feet — how many of each are there?" and it may jump straight to a conclusion, with no way to backtrack when it's wrong. A reasoning model instead generates a "thinking process" before the final answer:
User: There are 35 heads and 94 feet in the cage. How many chickens and how many rabbits?
(Chain of thought generated internally by the model, visible in reasoning models)
I have 35 heads; if all were chickens, there would be 35×2=70 feet.
There are actually 94 feet, so 94−70=24 feet are left over.
Each time a chicken is swapped for a rabbit, the head count stays the same and the feet increase by 2.
So rabbits = 24÷2 = 12, and chickens = 35−12 = 23.
Check: 12×4 + 23×2 = 48+46 = 94 ✓
Final answer: 23 chickens, 12 rabbits.This chain of thought can break the problem apart on its own, verify itself, and catch and fix its own mistakes, which makes its reasoning dramatically stronger than "answer in one shot." The behavior is internalized during training (see the next section on R1's key techniques) rather than coaxed out on the fly by user prompts.
2.2 Reasoning Models vs. Ordinary LLMs
| Dimension | Ordinary LLM (e.g., GPT-4o, DeepSeek-V3) | Reasoning model (e.g., o1, DeepSeek-R1) |
|---|---|---|
| Answer style | Generates the answer directly | Produces a chain of thought first, then the answer |
| Reasoning style | "Fast thinking" (System 1) | "Slow thinking" (System 2), with self-correction |
| Token overhead at inference | Low, low latency | High, high latency and cost |
| Hard math / code problems | Upper-middle; error-prone on complex steps | Much stronger; reliable over long reasoning chains |
| Simple tasks | Fast and accurate | Can be overkill — a sledgehammer to crack a nut |
| Relation to CoT prompting | Fires only when the user prompts "think step by step" | Internalized during training; produces CoT automatically |
Keep two levels distinct: the chain-of-thought (CoT) prompting covered in Prompt Engineering asks the model to think step by step at inference time — guidance at the prompt level — whereas a reasoning model hard-wires "long chains of thought" into model behavior at the training level, with no extra prompting needed. The former is memorizing the moves; the latter is building muscle memory.
Rule of thumb
Does your task need multi-step reasoning, self-correction, or code/math verification? If yes, reach for a reasoning model; if it's just Q&A, summarization, or rewriting, an ordinary LLM is faster and cheaper. Route with a cheaper model first — don't send every request down the "slow thinking" path.
2.3 Flagships of the Two Routes
- The closed-source standard-bearer: OpenAI's o1 / o3 series. o1 created the reasoning-model category and o3 pushed it further; OpenAI deliberately hides the full chain of thought and shows users only a summary, citing safety and monitorability;
- The open-source standard-bearer: DeepSeek-R1. Fully open weights (MIT license), a completely visible chain of thought, and self-hostable deployment — it has become the standard dissection specimen for reasoning-model research.
For a systematic view of the model-family landscape and the open/closed split, see Model and Leaderboard Quick Reference; to understand R1's architectural foundation, start with Transformers and Attention.
3. Inside R1's Technology: Reinforcement Learning, Distillation, and MoE
R1's technical approach boils down to three moves: use large-scale reinforcement learning to train reasoning in, use distillation to pass that reasoning ability down to smaller models, and use MoE to push runtime costs down.
3.1 Large-Scale Reinforcement Learning (RL) Elicits Reasoning
This is R1's most fundamental contribution. RLHF in traditional alignment (see Alignment: RLHF and DPO) aims to make models obedient; R1 aims reinforcement learning squarely at reasoning ability:
- What is trained: starting from DeepSeek-V3-Base, a cold-start fine-tune on a few thousand high-quality CoT samples, followed by large-scale RL;
- Reward signals: instead of an expensive reward model, it uses verifiable rules — whether the math answer is correct, whether the code passes its tests (checkable rewards);
- Algorithm: GRPO (Group Relative Policy Optimization), which estimates advantages through relative comparisons within a group of samples generated for the same question, eliminating the separate critic value model and slashing memory and compute requirements;
- Two-stage evolution: first came R1-Zero, built without the cold start (pure RL, from which spontaneous reasoning and "reflection and self-correction" behaviors emerged); cold-start data and human-preference alignment were then layered on to produce R1, balancing capability with readability.
R1 training pipeline (simplified)
① V3-Base → cold-start fine-tuning (a small set of high-quality CoT samples)
② Large-scale RL (rule-based rewards: math answer correctness / code passing tests) → R1-Zero-level reasoning
③ Rejection sampling to collect new data → SFT → RL again (adding human-preference alignment) → final R1The takeaway in one sentence: reasoning ability can be trained, and the reward signal doesn't have to come from human scoring — verifiable rules are enough.
3.2 Distillation: Passing a Large Model's Reasoning to Smaller Models
Using reasoning data generated by R1, the team fine-tuned a series of smaller open-source models, producing the R1-Distill series:
| Distilled model | Base | Parameter count | Highlights |
|---|---|---|---|
| R1-Distill-Qwen-1.5B | Qwen2.5-Math | 1.5B | Ultra-light, good for demos |
| R1-Distill-Qwen-7B / 14B / 32B | Qwen2.5 | 7B/14B/32B | The cost-effectiveness sweet spot |
| R1-Distill-Llama-8B / 70B | Llama 3.1/3.3 | 8B/70B | Mature ecosystem, widely deployed |
According to the technical report, R1-Distill-Qwen-32B scores 72.6% pass@1 on the AIME 2024 math competition, beating OpenAI's o1-mini — a 32B model matching closed-source reasoning models purely through what distillation passed down. For the principles and caveats behind distillation, see Fine-Tuning and PEFT (LoRA): distillation transfers a "behavior distribution" — the small model learns "how the large model thinks," not "everything the large model knows."
3.3 MoE Architecture: 671B Parameters of Capability, 37B of Compute
R1 shares its base with DeepSeek-V3 and uses a Mixture-of-Experts (MoE) architecture:
| Item | Value | Meaning |
|---|---|---|
| Total parameters | 671B | Large knowledge capacity |
| Activated parameters | ~37B | Only about 1/18 of the parameters are used per inference |
| Context length | 128K tokens | Supports long documents and long chains of thought |
| License | MIT (open weights) | Commercial use and derivative work allowed |
MoE breaks the equivalence of "large model" with "expensive model": total parameters set the capability ceiling, while activated parameters determine the cost per query. Combined with quantization, speculative sampling, and other inference optimization techniques, R1 can run on consumer GPUs (a quantized 4090, say) or a single A100 — unthinkable for closed models of the same generation. For complete deployment walkthroughs, see Deployment and Inference Optimization in Practice.
4. The Evolution of Reasoning Models
R1 is not an isolated case but a key milestone on the reasoning-model timeline:
2024-09 o1 (OpenAI, closed) ── defines the category: chain-of-thought + RL at inference time
2024-12 o3 (OpenAI, closed) ── stronger code/math, agent tasks
2025-01 DeepSeek-R1 (open source) ── same-tier capability, sharply lower cost, open weights
2025-02 DeepSeek-R1-V (multimodal reasoning) ── reasoning expands from pure text to text + images
…later OpenAI o3-mini / GPT-5 series, Claude thinking, Gemini reasoning modesReasoning Model Family Comparison (as of dataAsOf 2025-06)
| Model | Released | Open source | Base architecture | Notes |
|---|---|---|---|---|
| OpenAI o1 | 2024-09 | No | Proprietary (rumored GPT-4 family base) | Category creator, strong at math/code |
| OpenAI o3 | 2024-12 | No | Proprietary | Enhanced adaptive inference-time compute; a big lead on agent benchmarks |
| DeepSeek-R1 | 2025-01 | Yes (MIT) | V3 (MoE, 671B / 37B activated) | Matches o1, sharply lower cost, fully visible CoT |
| DeepSeek-R1-V | 2025-02 | Yes | V3 + vision encoder | First open-source multimodal reasoning model |
| Claude (thinking mode) | 2025 | No | Proprietary | Thinking mode can be toggled on and off |
| Gemini reasoning modes | 2025 | Partially | Proprietary | Deeply integrated into Gemini products |
R1 on Benchmarks (per the Technical Report)
| Benchmark | R1 | o1 | Notes |
|---|---|---|---|
| AIME 2024 (math competition) | 79.8% pass@1 | 79.2% | Tied |
| MATH-500 | 97.3% | 96.4% | Near the ceiling |
| Codeforces (competitive programming) | 2029 Elo | ~2027 Elo | Both expert-level |
| SWE-bench Verified | 49.2% | 48.9% (o1) | Real-world software engineering |
The numbers alone are not crushing; the point is doing the same thing at a hundredth of the cost. That is what detonated the chain reaction that followed.
5. Impact on the Industry
5.1 Falling Inference Costs → A Much Lower Bar for Agent Applications
What an agent needs most is planning, error correction, and multi-step execution — exactly what reasoning models are good at. Previously, handing an agent's "thinking" to o1 meant high costs and restricted API access; once R1 was open-sourced, teams could self-host a reasoning model as their agent's brain and bring costs into a manageable range.
- Planning: R1's CoT is naturally suited to task decomposition (what to do first, what comes next, how to verify);
- Tool calling: R1 performs well on Tool Use benchmarks and can hook up code interpreters, search, and other external tools;
- Inference-time compute: the "thinking length" is controllable (long thinking for hard tasks, short for easy ones), so cost and quality are tunable.
Reasoning models have therefore become a major driving force behind AI agent adoption; for the engineering practice, see Building an Agent from Scratch, and for product shapes, see Manus and Agent Applications. A handy way to put it: o1 proved the agent "brain" was feasible; R1 made that brain cheap enough for every team to afford.
Don't Treat Reasoning Models as the Default
A reasoning model's strength is concentrated in tasks that require multi-step reasoning. For conversation, writing, and knowledge Q&A, it is often slower, pricier, and no more accurate. Before shipping, run an evaluation round comparing ordinary and reasoning models, and let the data decide routing rather than gut feel. See Building an LLM Evaluation Suite for evaluation methods.
5.2 Open-Source Models Close In on Closed Models
Before R1, the gap between the open-source community and closed-source flagships was widely seen as "more than a generation"; after R1, it compressed to "the same generation":
| Dimension | Open-source models before R1 | After R1 |
|---|---|---|
| Reasoning | Clearly behind (no chain-of-thought training) | Math/code at closed-flagship level |
| Usability | Could chat and write, but struggled to reason | Deployable, commercially usable, visible CoT |
| Ecosystem | Research-oriented | Broad support from cloud providers and consumer hardware |
For developers, this means that beyond "renting capability through an API" there is now a second path: "deploying open-source models" — new options for both data security and cost control. For the full open/closed landscape, see Model and Leaderboard Quick Reference.
5.3 "Reinforcement Learning + Reasoning" Becomes a New Training Paradigm
R1 shifted the focus of post-training from RLHF's "match human preferences" to RL's "amplify capability" — training thinking with verifiable rule-based rewards. The paradigm was rapidly copied across the industry: the Qwen team distilled QwQ along similar lines, Gemini and Claude launched reasoning modes, and multiple startups went heavily into RL post-training. Why does the "training paradigm" matter more than any single model? Because it means capability gains are starting to decouple, in part, from dependence on massive-scale pretraining — teams short on compute can use "smarter training algorithms" to leapfrog. For frontier trends, see Frontier Research; for how to read the classic papers, see Paper Map.
6. Limitations and Controversies
R1's buzz doesn't mean it is without weaknesses. Here are the issues you can't get around:
6.1 Latency and Cost from Long Chains of Thought
A reasoning model "thinks for a long time" before every answer — a hard math problem can produce a chain of thought thousands or even tens of thousands of tokens long:
| Scenario | Ordinary LLM | Reasoning model |
|---|---|---|
| Simple Q&A | Fast, about 1 second | Slow, possibly 10+ seconds |
| Complex reasoning | Fast but may be wrong | Slow but more accurate |
| Cost (per token) | Low | High (several to over ten times more tokens) |
Practical rule: don't use reasoning models for simple tasks. A common pattern is "routing" — a cheap model first judges task difficulty, and only hard tasks go to a reasoning model; it is the classic pattern discussed on the Inference Optimization and Quantization page.
6.2 The Chain-of-Thought Verifiability Controversy
- The closed-source camp: OpenAI declines to publish o1's full chain of thought, citing monitorability and safety (preventing misuse and jailbreak steering), but this leaves the "reasoning process" unauditable by users;
- The open-source camp: R1's chain of thought is fully visible, yet some question whether "a visible CoT necessarily reflects the real reasoning" — the model may "perform thinking" to please the reward, the so-called reward hacking risk;
- The distillation controversy: a large share of R1's training data is synthetic data from the GPT-4 family (as the technical report acknowledges), fueling debate about "open-source models depending on data produced by closed models."
This trade-off among capability, transparency, and safety is exactly the question AI Safety and Governance set out to answer.
6.3 Benchmark Saturation
R1 hits 97.3% on MATH-500 and nearly 80% on AIME 2024 — numbers approaching the point where going higher no longer discriminates:
- Benchmark saturation: math-competition benchmarks were quickly exhausted, and differences between models are hard to separate with old benchmarks;
- Evaluation migration: the industry is moving to harder benchmarks that track real work more closely (SWE-bench, FrontierMath, Humanity's Last Exam, and others);
- Lesson: don't judge models by leaderboard scores alone; evaluate holistically against task type and cost. For methodology, see LLM Evaluation and Benchmarks; for hands-on practice, see Building an LLM Evaluation Suite.
7. Lessons from the Case
The R1 case leaves practitioners three core lessons:
First, reasoning ability can be trained, not just bought with scale. Using reinforcement learning with rule-based rewards, R1 lifted reasoning a full notch without a major increase in pretraining investment. Flipped around, many teams' pain point may not be "the model isn't big enough" but "the post-training isn't done right" — check the alignment and RL pipeline before reaching for a bigger model.
Second, the value of the open-source ecosystem is severely underestimated. R1's distilled series lets small and mid-sized teams acquire reasoning ability at very low cost, and the open weights in turn accelerated the inference-optimization ecosystem (quantization, frameworks, and deployment tools all moved quickly). For budget-constrained teams, open-source models + fine-tuning is often more controllable and cheaper than APIs.
Third, the "compute-for-cost" trade has been validated. R1 swapped "more thinking tokens + smarter training" for "less upfront investment," and this shift of cost from the training phase to the inference phase is reshaping LLM business models — for end users, an affordable strong-reasoning model makes many previously uneconomical applications (agents, coding assistants, data analysis) viable.
To return to the opening claim: R1's significance lies not in how many benchmarks it topped, but in proving that "high-performance reasoning" is no longer a privilege only money-burning giants can afford. To see how this shift extends into conversational products and coding tools, read ChatGPT and Conversational AI and GitHub Copilot and Code Intelligence; to trace the full technical storyline, start with A Brief History of AI and What Is AI? Hot Concepts.
Further Reading
- Large Language Models (LLMs) — the foundation of reasoning models: pretraining, instruction tuning, and alignment
- Prompt Engineering — CoT prompting: the "prompt-layer" ancestor of reasoning models
- Alignment: RLHF and DPO — understanding R1's reinforcement learning and human-preference alignment
- Fine-Tuning and PEFT (LoRA) — the principles and practice of distilling small models
- AI Agents — how reasoning models lower the barrier to agent deployment
- LLM Evaluation and Benchmarks — benchmark saturation and the evolution of new benchmarks
- Model and Leaderboard Quick Reference — side-by-side comparison of open and closed reasoning models
- Frontier Research — the latest work on the "reinforcement learning + reasoning" paradigm
- Deployment and Inference Optimization in Practice — quantizing and deploying MoE models
References
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025, arXiv:2501.12948) — R1's official technical report, with training details and benchmark data
- The official DeepSeek-R1 GitHub repository — weight downloads, license, and model card
- OpenAI. Introducing OpenAI o1 (official blog, 2024-09) — the official announcement that founded the reasoning-model category
- OpenAI. OpenAI o3 and o4-mini (official blog, 2024-12) — the o3 series launch and adaptive inference-time compute
- Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022) — the foundational paper on chain-of-thought prompting
- OpenAI. Learning to Reason with LLMs (official blog, 2024-09) — the official explanation of o1's training method (RL at inference time)