Skip to content

DeepSeek-R1 and Reasoning Models

At a glance In January 2025, DeepSeek-R1 took on OpenAI's o1 as an open-source model and turned "reasoning models" from a lab term into a global hotspot; this article breaks down the chain-of-thought mechanism, R1's reinforcement learning and distillation techniques, the evolution of reasoning models, and their industry impact, limitations, and controversies.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

DeepSeek-R1 and Reasoning Models ​

A reasoning model is a large language model (LLM) that produces an internal chain-of-thought (CoT) before generating its final answer — think it through first, then respond. Rather than blurting out a conclusion the way an ordinary LLM does, it writes its "thinking process" explicitly into the generation sequence. In January 2025, DeepSeek released and open-sourced the reasoning model DeepSeek-R1, which — at a training cost reportedly far below that of comparable models — matched or even beat OpenAI's o1 on reasoning tasks such as math and code. The release turned the "reasoning model" from a laboratory term into a global talking point and reshaped the fundamentals of open versus closed models, cost versus capability, training paradigms, and agent deployment.

Using R1 as our lens, this article makes three things clear: what actually separates reasoning models from ordinary LLMs, which exceptionally cost-effective techniques R1 employed, and the ripple effects it triggered across the industry. If you want to brush up on the basics first, read alongside Large Language Models (LLMs) and Prompt Engineering.

1. Background: A Global "Low-Cost, High-Performance" Phenomenon ​

1.1 Timeline and How It Unfolded ​

DateEventSignificance
2024-09OpenAI releases o1-preview / o1-miniThe "reasoning model" category is born, showcasing chain-of-thought reasoning
2024-12OpenAI releases o3 (with an o3-mini preview)Reasoning takes another step up, but remains closed-source
2025-01-20DeepSeek-R1 officially released and open-sourced (MIT-licensed weights)o1-level reasoning, openly available
2025-02DeepSeek-R1-V (multimodal reasoning) and a 167B MoE reasoning variant releasedReasoning extends to multimodality and broader deployment

The day R1 launched, it ignited a global media storm: downloads of its weights on Hugging Face surged within hours, multiple cloud providers (Nvidia, Microsoft, Amazon, and others) announced support, and the app briefly topped the App Store free charts. For an open-source model from a Chinese startup, that level of attention was unprecedented — the industry widely dubbed it "AI's Sputnik moment."

1.2 The DeepSeek Story and the "Low-Cost, High-Performance" Narrative ​

What made R1 stunning was not that it was the most capable model, but its price-to-performance ratio. According to the technical report and industry estimates, R1's training cost was far below what OpenAI's comparable models are rumored to have cost:

  • RL-stage compute: the report states that R1-Zero's reinforcement learning training used only about 8,000 V100 GPU hours — typically the budget for a small experiment at a big tech company;
  • Overall training cost estimate: taking DeepSeek-V3's reported ~2.78 million H800 GPU hours and roughly US$5.57 million (industry estimate) as a reference, R1's full training cost is widely believed to be in the low millions of dollars — versus rumored training budgets of hundreds of millions of dollars for closed models, a gap of two to three orders of magnitude;
  • Inference and deployment costs: R1 uses a Mixture-of-Experts (MoE) architecture with 671B total parameters but only about 37B activated per inference, and with the quantization and inference optimizations the open-source community adapted quickly (see Inference Optimization and Quantization), the cost per query was driven extremely low.

"Algorithmic innovation can partially make up for a compute gap" — before R1 this was a hypothesis; after R1 it became industry consensus. It also explains why the DeepSeek family has been a permanent fixture on the Model and Leaderboard Quick Reference page ever since.

2. What Is a Reasoning Model ​

2.1 The Mechanism: Making "Thinking" Part of Generation ​

An ordinary LLM works by input → answer, directly: ask it the classic puzzle "Chickens and rabbits are shut in a cage; together they have 35 heads and 94 feet — how many of each are there?" and it may jump straight to a conclusion, with no way to backtrack when it's wrong. A reasoning model instead generates a "thinking process" before the final answer:

User: There are 35 heads and 94 feet in the cage. How many chickens and how many rabbits?

(Chain of thought generated internally by the model, visible in reasoning models)
I have 35 heads; if all were chickens, there would be 35×2=70 feet.
There are actually 94 feet, so 94−70=24 feet are left over.
Each time a chicken is swapped for a rabbit, the head count stays the same and the feet increase by 2.
So rabbits = 24÷2 = 12, and chickens = 35−12 = 23.
Check: 12×4 + 23×2 = 48+46 = 94 ✓

Final answer: 23 chickens, 12 rabbits.

This chain of thought can break the problem apart on its own, verify itself, and catch and fix its own mistakes, which makes its reasoning dramatically stronger than "answer in one shot." The behavior is internalized during training (see the next section on R1's key techniques) rather than coaxed out on the fly by user prompts.

2.2 Reasoning Models vs. Ordinary LLMs ​

DimensionOrdinary LLM (e.g., GPT-4o, DeepSeek-V3)Reasoning model (e.g., o1, DeepSeek-R1)
Answer styleGenerates the answer directlyProduces a chain of thought first, then the answer
Reasoning style"Fast thinking" (System 1)"Slow thinking" (System 2), with self-correction
Token overhead at inferenceLow, low latencyHigh, high latency and cost
Hard math / code problemsUpper-middle; error-prone on complex stepsMuch stronger; reliable over long reasoning chains
Simple tasksFast and accurateCan be overkill — a sledgehammer to crack a nut
Relation to CoT promptingFires only when the user prompts "think step by step"Internalized during training; produces CoT automatically

Keep two levels distinct: the chain-of-thought (CoT) prompting covered in Prompt Engineering asks the model to think step by step at inference time — guidance at the prompt level — whereas a reasoning model hard-wires "long chains of thought" into model behavior at the training level, with no extra prompting needed. The former is memorizing the moves; the latter is building muscle memory.

Rule of thumb

Does your task need multi-step reasoning, self-correction, or code/math verification? If yes, reach for a reasoning model; if it's just Q&A, summarization, or rewriting, an ordinary LLM is faster and cheaper. Route with a cheaper model first — don't send every request down the "slow thinking" path.

2.3 Flagships of the Two Routes ​

  • The closed-source standard-bearer: OpenAI's o1 / o3 series. o1 created the reasoning-model category and o3 pushed it further; OpenAI deliberately hides the full chain of thought and shows users only a summary, citing safety and monitorability;
  • The open-source standard-bearer: DeepSeek-R1. Fully open weights (MIT license), a completely visible chain of thought, and self-hostable deployment — it has become the standard dissection specimen for reasoning-model research.

For a systematic view of the model-family landscape and the open/closed split, see Model and Leaderboard Quick Reference; to understand R1's architectural foundation, start with Transformers and Attention.

3. Inside R1's Technology: Reinforcement Learning, Distillation, and MoE ​

R1's technical approach boils down to three moves: use large-scale reinforcement learning to train reasoning in, use distillation to pass that reasoning ability down to smaller models, and use MoE to push runtime costs down.

3.1 Large-Scale Reinforcement Learning (RL) Elicits Reasoning ​

This is R1's most fundamental contribution. RLHF in traditional alignment (see Alignment: RLHF and DPO) aims to make models obedient; R1 aims reinforcement learning squarely at reasoning ability:

  • What is trained: starting from DeepSeek-V3-Base, a cold-start fine-tune on a few thousand high-quality CoT samples, followed by large-scale RL;
  • Reward signals: instead of an expensive reward model, it uses verifiable rules — whether the math answer is correct, whether the code passes its tests (checkable rewards);
  • Algorithm: GRPO (Group Relative Policy Optimization), which estimates advantages through relative comparisons within a group of samples generated for the same question, eliminating the separate critic value model and slashing memory and compute requirements;
  • Two-stage evolution: first came R1-Zero, built without the cold start (pure RL, from which spontaneous reasoning and "reflection and self-correction" behaviors emerged); cold-start data and human-preference alignment were then layered on to produce R1, balancing capability with readability.
R1 training pipeline (simplified)
① V3-Base → cold-start fine-tuning (a small set of high-quality CoT samples)
② Large-scale RL (rule-based rewards: math answer correctness / code passing tests) → R1-Zero-level reasoning
③ Rejection sampling to collect new data → SFT → RL again (adding human-preference alignment) → final R1

The takeaway in one sentence: reasoning ability can be trained, and the reward signal doesn't have to come from human scoring — verifiable rules are enough.

3.2 Distillation: Passing a Large Model's Reasoning to Smaller Models ​

Using reasoning data generated by R1, the team fine-tuned a series of smaller open-source models, producing the R1-Distill series:

Distilled modelBaseParameter countHighlights
R1-Distill-Qwen-1.5BQwen2.5-Math1.5BUltra-light, good for demos
R1-Distill-Qwen-7B / 14B / 32BQwen2.57B/14B/32BThe cost-effectiveness sweet spot
R1-Distill-Llama-8B / 70BLlama 3.1/3.38B/70BMature ecosystem, widely deployed

According to the technical report, R1-Distill-Qwen-32B scores 72.6% pass@1 on the AIME 2024 math competition, beating OpenAI's o1-mini — a 32B model matching closed-source reasoning models purely through what distillation passed down. For the principles and caveats behind distillation, see Fine-Tuning and PEFT (LoRA): distillation transfers a "behavior distribution" — the small model learns "how the large model thinks," not "everything the large model knows."

3.3 MoE Architecture: 671B Parameters of Capability, 37B of Compute ​

R1 shares its base with DeepSeek-V3 and uses a Mixture-of-Experts (MoE) architecture:

ItemValueMeaning
Total parameters671BLarge knowledge capacity
Activated parameters~37BOnly about 1/18 of the parameters are used per inference
Context length128K tokensSupports long documents and long chains of thought
LicenseMIT (open weights)Commercial use and derivative work allowed

MoE breaks the equivalence of "large model" with "expensive model": total parameters set the capability ceiling, while activated parameters determine the cost per query. Combined with quantization, speculative sampling, and other inference optimization techniques, R1 can run on consumer GPUs (a quantized 4090, say) or a single A100 — unthinkable for closed models of the same generation. For complete deployment walkthroughs, see Deployment and Inference Optimization in Practice.

4. The Evolution of Reasoning Models ​

R1 is not an isolated case but a key milestone on the reasoning-model timeline:

2024-09  o1 (OpenAI, closed) ── defines the category: chain-of-thought + RL at inference time
2024-12  o3 (OpenAI, closed) ── stronger code/math, agent tasks
2025-01  DeepSeek-R1 (open source) ── same-tier capability, sharply lower cost, open weights
2025-02  DeepSeek-R1-V (multimodal reasoning) ── reasoning expands from pure text to text + images
…later   OpenAI o3-mini / GPT-5 series, Claude thinking, Gemini reasoning modes

Reasoning Model Family Comparison (as of dataAsOf 2025-06) ​

ModelReleasedOpen sourceBase architectureNotes
OpenAI o12024-09NoProprietary (rumored GPT-4 family base)Category creator, strong at math/code
OpenAI o32024-12NoProprietaryEnhanced adaptive inference-time compute; a big lead on agent benchmarks
DeepSeek-R12025-01Yes (MIT)V3 (MoE, 671B / 37B activated)Matches o1, sharply lower cost, fully visible CoT
DeepSeek-R1-V2025-02YesV3 + vision encoderFirst open-source multimodal reasoning model
Claude (thinking mode)2025NoProprietaryThinking mode can be toggled on and off
Gemini reasoning modes2025PartiallyProprietaryDeeply integrated into Gemini products

R1 on Benchmarks (per the Technical Report) ​

BenchmarkR1o1Notes
AIME 2024 (math competition)79.8% pass@179.2%Tied
MATH-50097.3%96.4%Near the ceiling
Codeforces (competitive programming)2029 Elo~2027 EloBoth expert-level
SWE-bench Verified49.2%48.9% (o1)Real-world software engineering

The numbers alone are not crushing; the point is doing the same thing at a hundredth of the cost. That is what detonated the chain reaction that followed.

5. Impact on the Industry ​

5.1 Falling Inference Costs → A Much Lower Bar for Agent Applications ​

What an agent needs most is planning, error correction, and multi-step execution — exactly what reasoning models are good at. Previously, handing an agent's "thinking" to o1 meant high costs and restricted API access; once R1 was open-sourced, teams could self-host a reasoning model as their agent's brain and bring costs into a manageable range.

  • Planning: R1's CoT is naturally suited to task decomposition (what to do first, what comes next, how to verify);
  • Tool calling: R1 performs well on Tool Use benchmarks and can hook up code interpreters, search, and other external tools;
  • Inference-time compute: the "thinking length" is controllable (long thinking for hard tasks, short for easy ones), so cost and quality are tunable.

Reasoning models have therefore become a major driving force behind AI agent adoption; for the engineering practice, see Building an Agent from Scratch, and for product shapes, see Manus and Agent Applications. A handy way to put it: o1 proved the agent "brain" was feasible; R1 made that brain cheap enough for every team to afford.

Don't Treat Reasoning Models as the Default

A reasoning model's strength is concentrated in tasks that require multi-step reasoning. For conversation, writing, and knowledge Q&A, it is often slower, pricier, and no more accurate. Before shipping, run an evaluation round comparing ordinary and reasoning models, and let the data decide routing rather than gut feel. See Building an LLM Evaluation Suite for evaluation methods.

5.2 Open-Source Models Close In on Closed Models ​

Before R1, the gap between the open-source community and closed-source flagships was widely seen as "more than a generation"; after R1, it compressed to "the same generation":

DimensionOpen-source models before R1After R1
ReasoningClearly behind (no chain-of-thought training)Math/code at closed-flagship level
UsabilityCould chat and write, but struggled to reasonDeployable, commercially usable, visible CoT
EcosystemResearch-orientedBroad support from cloud providers and consumer hardware

For developers, this means that beyond "renting capability through an API" there is now a second path: "deploying open-source models" — new options for both data security and cost control. For the full open/closed landscape, see Model and Leaderboard Quick Reference.

5.3 "Reinforcement Learning + Reasoning" Becomes a New Training Paradigm ​

R1 shifted the focus of post-training from RLHF's "match human preferences" to RL's "amplify capability" — training thinking with verifiable rule-based rewards. The paradigm was rapidly copied across the industry: the Qwen team distilled QwQ along similar lines, Gemini and Claude launched reasoning modes, and multiple startups went heavily into RL post-training. Why does the "training paradigm" matter more than any single model? Because it means capability gains are starting to decouple, in part, from dependence on massive-scale pretraining — teams short on compute can use "smarter training algorithms" to leapfrog. For frontier trends, see Frontier Research; for how to read the classic papers, see Paper Map.

6. Limitations and Controversies ​

R1's buzz doesn't mean it is without weaknesses. Here are the issues you can't get around:

6.1 Latency and Cost from Long Chains of Thought ​

A reasoning model "thinks for a long time" before every answer — a hard math problem can produce a chain of thought thousands or even tens of thousands of tokens long:

ScenarioOrdinary LLMReasoning model
Simple Q&AFast, about 1 secondSlow, possibly 10+ seconds
Complex reasoningFast but may be wrongSlow but more accurate
Cost (per token)LowHigh (several to over ten times more tokens)

Practical rule: don't use reasoning models for simple tasks. A common pattern is "routing" — a cheap model first judges task difficulty, and only hard tasks go to a reasoning model; it is the classic pattern discussed on the Inference Optimization and Quantization page.

6.2 The Chain-of-Thought Verifiability Controversy ​

  • The closed-source camp: OpenAI declines to publish o1's full chain of thought, citing monitorability and safety (preventing misuse and jailbreak steering), but this leaves the "reasoning process" unauditable by users;
  • The open-source camp: R1's chain of thought is fully visible, yet some question whether "a visible CoT necessarily reflects the real reasoning" — the model may "perform thinking" to please the reward, the so-called reward hacking risk;
  • The distillation controversy: a large share of R1's training data is synthetic data from the GPT-4 family (as the technical report acknowledges), fueling debate about "open-source models depending on data produced by closed models."

This trade-off among capability, transparency, and safety is exactly the question AI Safety and Governance set out to answer.

6.3 Benchmark Saturation ​

R1 hits 97.3% on MATH-500 and nearly 80% on AIME 2024 — numbers approaching the point where going higher no longer discriminates:

  • Benchmark saturation: math-competition benchmarks were quickly exhausted, and differences between models are hard to separate with old benchmarks;
  • Evaluation migration: the industry is moving to harder benchmarks that track real work more closely (SWE-bench, FrontierMath, Humanity's Last Exam, and others);
  • Lesson: don't judge models by leaderboard scores alone; evaluate holistically against task type and cost. For methodology, see LLM Evaluation and Benchmarks; for hands-on practice, see Building an LLM Evaluation Suite.

7. Lessons from the Case ​

The R1 case leaves practitioners three core lessons:

First, reasoning ability can be trained, not just bought with scale. Using reinforcement learning with rule-based rewards, R1 lifted reasoning a full notch without a major increase in pretraining investment. Flipped around, many teams' pain point may not be "the model isn't big enough" but "the post-training isn't done right" — check the alignment and RL pipeline before reaching for a bigger model.

Second, the value of the open-source ecosystem is severely underestimated. R1's distilled series lets small and mid-sized teams acquire reasoning ability at very low cost, and the open weights in turn accelerated the inference-optimization ecosystem (quantization, frameworks, and deployment tools all moved quickly). For budget-constrained teams, open-source models + fine-tuning is often more controllable and cheaper than APIs.

Third, the "compute-for-cost" trade has been validated. R1 swapped "more thinking tokens + smarter training" for "less upfront investment," and this shift of cost from the training phase to the inference phase is reshaping LLM business models — for end users, an affordable strong-reasoning model makes many previously uneconomical applications (agents, coding assistants, data analysis) viable.

To return to the opening claim: R1's significance lies not in how many benchmarks it topped, but in proving that "high-performance reasoning" is no longer a privilege only money-burning giants can afford. To see how this shift extends into conversational products and coding tools, read ChatGPT and Conversational AI and GitHub Copilot and Code Intelligence; to trace the full technical storyline, start with A Brief History of AI and What Is AI? Hot Concepts.

Further Reading ​

References ​