Skip to content

Frontier Trends

At a glance A 2023–2025 LLM frontier trends overview: inference-time scaling (o1), long context and million-token windows, MoE at massive scale, native multimodal, Agents and tool use, Mamba state space models, quantization and efficient inference, open-source catching up to closed-source, and new evaluation paradigms — each with evolution timeline, representative papers/models, and a one-line takeaway.

This page contains time-sensitive content, current as of 2025-06; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

Frontier Trends ​

One-sentence summary: This page is a "survey of nine major LLM frontier trends from 2023–2025" — each trend includes its evolution timeline, representative papers/models, what happened, and a one-line takeaway. It doesn't aim to be exhaustive; instead, after reading it, you should be able to answer "what's been happening this half-year, what's the story behind each trend, and what does it mean for my work?"

1. How to Read "Frontier" ​

  1. This is the "frontier" row of the map expanded: Pair with Paper Map. Content is based on publicly available information through mid-2025; time-sensitive content should be verified against official releases.
  2. Frontier information has a short half-life: Numbers (benchmark scores, parameter counts) go stale in months; trends (how paradigms evolve) last at least two to three years. This page helps you capture the latter.
  3. The judgment criterion is "whether a paradigm is established": Is a paper "a single experiment result" or "the starting point of a new paradigm"? For example, FlashAttention is a paradigm (I/O-aware), while a new attention variant might just be one experiment. Apply the paper quality judgment criteria and you'll save a lot of time.
  4. Methods for staying current are in the "following arXiv" section of Reading Discipline & FAQ.

1. Inference-Time Scaling: From "Pretraining Scale" to "Inference Compute" (o1, test-time compute) ​

Evolution timeline: The idea of test-time compute isn't new — AlphaGo's Monte Carlo Tree Search (2016) was essentially "trade more inference compute for a better answer." But on language models, it was long overshadowed by "pretraining scale." In September 2024, OpenAI released o1, the first to productize "reinforcement learning + chain of thought + inference-time search/reflection" — the model "thinks" before answering, and the more it thinks (more test-time compute), the higher the accuracy. o3 (announced late 2024) follows the same route. On the open-source side, DeepSeek-R1 (January 2025) reproduced "reasoning emergence" using pure reinforcement learning (GRPO), and released a full technical report and weights.

  • Representative papers/models: OpenAI o1 (official blog, no arXiv), DeepSeek-R1 (2025).
  • One-line takeaway: This is the fourth scalable dimension after pretraining scale — an "inference-time scaling law" is taking shape, extending Scaling Laws from "training compute" to "inference compute," and changing both cost structure (inference becomes more expensive) and evaluation methodology (with-thinking vs. without-thinking must be measured separately).
  • Implications for practitioners: If your work involves high-value math, code, or logic problems, "whether to use a reasoning model" becomes a cost decision: more thinking → higher accuracy, but the cost per response may be an order of magnitude higher. When evaluating, always separate "thinking" and "non-thinking" modes — otherwise conclusions will be distorted.

Why it matters

Before 2022, "stronger = bigger pretraining." o1 proved that the same weights can trade "thinking budget" for capability. The product implication is a direct economic calculation: should you spend 10× more inference cost for one response in exchange for higher accuracy? This tradeoff will become a daily decision for every LLM application.

2. Long Context and Million-Token Windows ​

Evolution timeline: Context windows grew from GPT-3's 2K all the way to GPT-4 Turbo's 128K, Claude's 200K, and Gemini 1.5's million-token (1M) window. The technical foundation has three layers: positional encoding extrapolation (ALiBi, YaRN, NTK-aware), efficient attention (FlashAttention series), and long-context training data upsampling. After 2024, "needle-in-a-haystack" evaluations showed models can "remember" key information in long text, but engineering challenges (cost, retrieval, long-range fact extraction) remain.

  • Representative papers/models: Gemini 1.5 Technical Report (2024), YaRN (2023), LongLoRA (2023), FlashAttention (2022).
  • One-line takeaway: "Fitting it in" is solved; "using it effectively" is the new battlefield — long context and RAG are complementary, not substitutes: the former suits "read an entire document," the latter suits "needle-in-a-haystack" (see Context and Long Context).
  • Implications for practitioners: When choosing a solution, first consider the task type: use long context for analyzing a whole book or contract, use RAG for knowledge-base QA. Also remember that long context means attention and KV Cache costs scale linearly with length — it's not a free lunch.

3. MoE at Massive Scale: Sparse Activation Becomes the Main Cost-Effective Choice ​

Evolution timeline: MoE went from lab demos at GShard (2020) and Switch Transformer (2021), through GLM proving "same compute, stronger," to Mixtral 8×7B (December 2023, 46.7B total / ~12.9B activated) bringing sparsity to open-source. DeepSeek-V3 (December 2024) with 671B total / 37B activated parameters and ~1/10 the training cost caught up to equivalent dense models; GPT-4 is also rumored to use an ~8-expert architecture. In 2025, Qwen3-MoE, Llama 4, and others pushed MoE further into the mainstream.

  • Representative papers/models: Mixtral of Experts (2024), DeepSeek-V3 (2024), Switch Transformers (2021).
  • One-line takeaway: MoE rewrites "total params = capability ceiling" to "activated params = compute cost, total params = knowledge capacity" — it's the cost-effective standard answer at large parameter counts. The tradeoff is GPU memory (all experts must reside) and deployment complexity (see MoE and Super-Scale Models and MoE Sparse Experts).
  • Implications for practitioners: When choosing a base model, look at "activated parameters," not "total parameters." Before deploying MoE, calculate the memory budget (all experts resident + KV Cache) — you may need expert parallelism or CPU offload.

4. Native Multimodal: From "Patching" to "Native" ​

Evolution timeline: CLIP (2021) first solved the "image-text alignment" problem, Flamingo/LLaVA took a "vision encoder + projection + LLM" patching route; GPT-4V (2023) integrated vision into conversation. In 2024, GPT-4o and Gemini moved toward native multimodal (one model handling text/image/audio, real-time interaction), Sora put "video generation as world simulator" in the headlines, and open-source models like Qwen2.5-VL and InternVL caught up quickly, with multimodal evaluation (like MMMU) becoming an independent track.

  • Representative papers/models: GPT-4o (official blog, no arXiv), Gemini 1.5 (2024), Sora Technical Report (2024), CLIP (2021).
  • One-line takeaway: Modality boundaries are disappearing: unified tokenization + unified training objectives is the direction — "native" is currently more about architecture and interaction, and its alignment and safety complexity rises in tandem (see Multimodal LLMs).
  • Implications for practitioners: When your product needs image/voice input, prefer native multimodal APIs — don't patch CLIP+LLM yourself. The latter's debugging cost for multi-turn conversation and alignment is far higher than for text-only scenarios.

5. Agents and Tool Use: From "Conversation" to "Getting Things Done" ​

Evolution timeline: ReAct (2022) established the "reasoning + acting interleaved" paradigm; Toolformer (2023) let models learn to self-select tools; in 2023, function calling became an API standard, and AutoGPT proved "LLM loops calling tools" is feasible but unreliable. In 2024, MCP (Model Context Protocol, November 2024) standardized "tools/memory/resources," Devin showcased an "AI software engineer" product form, and multi-agent collaboration (MetaGPT, etc.) entered engineering practice.

  • Representative papers/models: ReAct (2022), Toolformer (2023), MetaGPT (2023).
  • One-line takeaway: Agent bottlenecks have shifted from "can it call tools?" to "reliability, memory, evaluation, and cost control" — prompt injection and loss-of-control risks are hard gates for deployment (see Agents with LLMs).
  • Implications for practitioners: Before deploying an Agent, do three things: whitelist tools and minimize permissions, run prompt injection tests, set per-task cost caps. Evaluate using "task completion rate," not "conversation fluency."

6. Mamba and State Space Models: Challenging Attention's Complexity ​

Evolution timeline: State space models (SSM) began serious study with S4 (2021); Mamba (December 2023) used selective state space models (S6) to achieve linear-complexity sequence modeling that's also hardware-friendly, outperforming same-scale Transformers on long sequences. Subsequently, Mamba-2 (SSD unified attention and SSM), Jamba (Mamba+Transformer hybrid), Zamba, etc. emerged. Hybrid architecture — attention for retrieval/local, SSM for long-range compression — is the most pragmatic route.

  • Representative papers/models: Mamba (2023), Mamba-2 (2024), Jamba (2024).
  • One-line takeaway: "Pure attention" may not be the endgame; hybrid architectures are becoming the practical answer — but the strongest commercial/open-source models are still Transformer-based. "Replacement" is premature, but "hybrid" is already reality (see Transformer Architecture).
  • Implications for practitioners: For ultra-long-sequence + cost-sensitive scenarios, you can experiment with Mamba-family or hybrid models. But the ecosystem (fine-tuning, quantization, inference engines) still favors Transformer, so factor ecosystem maturity into your decision.

7. Quantization and Efficient Inference: Fitting Large Models on a Single GPU ​

Evolution timeline: Quantization evolved from post-training PTQ (GPTQ, AWQ) to 4-bit (GGUF/llama.cpp ecosystem) to FP8/FP4 training-and-inference unification. QLoRA made fine-tuning 65B on a single 24GB GPU possible. Inference engines (vLLM's PagedAttention, SGLang, TensorRT-LLM) boosted throughput by an order of magnitude. The combination of "sparse activation + quantization + continuous batching" has become the deployment standard.

  • Representative papers/models: QLoRA (2023), PagedAttention/vLLM (2023).
  • One-line takeaway: Quantization and batching have made "running open-source models locally" a routine operation — cost is no longer the primary barrier for the open-source route; accuracy loss is typically <1% for most tasks (see Deployment and Serving and Inference Fundamentals).
  • Implications for practitioners: The deployment standard flow is now the three-piece set: "quantization + continuous batching + inference engine." But the "<1% loss" claim only holds for routine tasks — sensitive scenarios (math, fact extraction) require pre/post quantization regression tests.

8. Open-Source Catching Up to Closed-Source ​

Evolution timeline: GPT-2 (2019) was open-source but small; Llama 1 (February 2023) opened the door; Llama 2 (July 2023) open-sourced weights + commercial license; Llama 3 (2024, 8B/70B/405B, 15T+ tokens) pushed the open-source ceiling to 405B. Mistral/Mixtral, Qwen series, DeepSeek-V3/R1, GLM, and others continued closing the gap with the GPT-4/Claude tier. From 2025 onward, open-source has matched or even surpassed closed-source on some leaderboards (math, code).

  • Representative papers/models: Llama 3 (2024), Llama 2 (2023), DeepSeek-V3 (2024).
  • One-line takeaway: The gap between open-source and closed-source is shifting from "capability" to "safety and productization" — open weights make audit, fine-tuning, and local deployment possible, but also introduce new issues around misuse and poisoning (see Llama and the Open-Source Ecosystem and Safety and Risks).
  • Implications for practitioners: Open-source makes "private deployment + domain fine-tuning" a viable route — it's the first choice when data must stay on-prem. But build supporting compliance and security review processes; don't treat "open-source" as "free."

9. New Evaluation Paradigms: From "Grinding Leaderboards" to "Measuring Real Capability" ​

Evolution timeline: After the GLUE/SuperGLUE era ended, MMLU/GSM8K/HumanEval became benchmarks — but they were quickly called out for saturation and data contamination. The community shifted toward: ① Human blind-aggregated evaluation (Chatbot Arena's Elo); ② LLM-as-a-judge (models score each other, but with position/self-preference biases); ③ Capability-stratified evaluation (math, code, long context, tool use measured separately); ④ New benchmarks tied to inference-time scaling (e.g., AIME, GPQA).

  • Representative papers/models: Chatbot Arena (2023, LMSYS), MT-Bench (2023).
  • One-line takeaway: "Single-benchmark scores" are losing meaning; "evaluation systems" have become an independent discipline for engineering and research — every SOTA claim should be probed on evaluation protocol and data contamination (see Evaluation and Benchmarks and Evals in Practice).
  • Implications for practitioners: Don't just follow leaderboards for your own evaluation: use the three-piece set of Arena-style human blind evaluation + LLM-as-a-judge + your own golden set, and rotate questions periodically to prevent contamination.
TrendRepresentative Papers/ModelsOne-Line TakeawaySite Coverage
Inference-time scalingo1, o3, DeepSeek-R1Fourth scalable dimension, changes cost and evaluation structureScaling Laws
Long contextGemini 1.5 (1M), YaRN, FlashAttentionFitting in ≠ using well; complementary with RAGContext and Long Context
MoE at scaleMixtral, DeepSeek-V3Activated params determine cost, sparsity is mainstreamMoE, MoE Cases
Native multimodalGPT-4o, Gemini, SoraModality boundaries disappearing, unified tokenizationMultimodal LLMs
Agents and toolsReAct, MCP, DevinBottlenecks in reliability/memory/evaluation/safetyAgents
Mamba / SSMMamba, Mamba-2, JambaLinear complexity feasible, hybrid architecture is pragmaticTransformer
Quantized inferenceQLoRA, vLLM, GPTQ/AWQLocal deployment becomes routine, cost is no longer the barrierDeployment
Open-source catching upLlama 3, Qwen, DeepSeekGap shifts to safety and productizationLlama Ecosystem
New evaluation paradigmsChatbot Arena, MT-BenchEvaluation as independent discipline, farewell to single scoresEvaluation

4. Frontier Timeline (2023–2025 Key Milestones) ​

Compressing the nine trends into a timeline, watching how they interweave:

TimeEventTrend Category
2023.03GPT-4 release (multimodal input, exam-level capability)Capability ceiling / Multimodal
2023.07Llama 2 open-source weights + commercial licenseOpen-source catching up
2023.09vLLM open-sourced (PagedAttention)Quantized inference / Systems
2023.12Mixtral 8×7B and Mamba released simultaneouslyMoE / SSM
2024.02Gemini 1.5 (100K tokens), Sora announcedLong context / Multimodal
2024.03Claude 3, Jamba (Mamba+Transformer hybrid)Open-source / Hybrid architecture
2024.05GPT-4o (native multimodal, real-time voice)Native multimodal
2024.07Llama 3 405B open-source (15T+ tokens)Open-source catching up
2024.09OpenAI o1 releasedInference-time scaling
2024.12DeepSeek-V3 (671B/37B activated, low-cost training)MoE at massive scale
2025.01DeepSeek-R1 (pure RL reasoning model, open-source)Inference-time scaling + Open-source
Mid-2025Open-source matches closed-source on math/code leaderboards (subject to official releases)Open-source catching up

Reading this timeline reveals a pattern: 2023's "model releases" became "systems + models dual releases" in 2024, and "model + method + weights fully open-source" in 2025 — openness and engineering complexity rise together, which is the force behind open-source catching up to closed-source.

The nine trends aren't isolated — they interlock:

RelationshipExplanation
Inference-time scaling ↔ Open-source catching upDeepSeek-R1 turned test-time from a closed-source direction into an open-source benchmark, mutually accelerating both
Long context ↔ Quantized inferenceFlashAttention made long context feasible; quantization and KV Cache optimization made long context costs bearable
MoE ↔ Open-source catching upDeepSeek-V3 used MoE to drive training costs down to ~1/10, a key chip in the open-source closing-the-gap story
Multimodal ↔ AgentsMultimodal agents (screen-reading, browser-operating) merge two trends into one product form
New evaluation paradigms ↔ All of the aboveEvery trend brings new evaluation needs, and evaluation in turn exposes new problems (e.g., evaluating o1's inference cost)

You don't need to track all nine — focus by your role:

RolePrioritySecondary
ML Engineer / Application DevQuantized inference, long context, Agents, open-source catching upInference-time scaling (API availability)
ResearcherInference-time scaling, Mamba/SSM, new evaluation paradigmsMoE at massive scale
Product / Non-MLOpen-source catching up, Agents, native multimodalOthers (awareness only)
Deployment / SREQuantized inference, MoE (memory and scheduling)Long context

7. How to Read a Frontier Paper (Example: DeepSeek-R1) ​

Frontier papers are read the same way as classic papers, just with an added "timeliness" variable: judge whether it's worth reading first, then decide how deep. Walking through DeepSeek-R1 as an example:

text
Got paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL

Pass 1 (15 min): Judge whether to deep-dive
├─ Abstract: Uses pure RL (GRPO) to incentivize reasoning, no SFT, open weights
├─ Figures: AIME/GPQA score comparisons, R1-Zero reflection behavior examples
├─ Conclusion: Inference-time scaling can be reproduced on open-source
└─ Judgment: Milestone-level (open weights + full tech report) — worth Pass 2

Pass 2 (1 hour): Understand the mechanism
├─ Method: GRPO replaces PPO (remove critic, use group-relative advantage), cold-start data + two-stage RL
├─ Ablations: What does RL contribute vs. RLVR (verifiable rewards)? Why is R1-Zero unstable?
├─ Limitations: Strongly dependent on tasks with "verifiable answers"; formatted rewards may overfit
└─ Deliverable: Can state in one sentence "why GRPO saves memory and why it's better suited for reasoning training than PPO"

Pass 3 (optional): Find improvement points
├─ Probe: Why is distillation to smaller models still effective? Can RL be replicated to code/Agent scenarios?
├─ Improvement hypothesis: Apply GRPO to tool-call training
└─ Deliverable: One-page related work + one verifiable hypothesis

This workflow is based on the principles in judging paper quality and three-pass reading: Pass 1 solves "should I read it," Pass 2 solves "how deep," Pass 3 solves "where to go next." The biggest pitfall with frontier papers is "deep-reading on the first pass" — spending two hours on a paper that wasn't worth it.

8. How to Stay Current ​

  1. Subscribe to arXiv abstracts (cs.CL + cs.LG), scan headlines weekly (method in Reading Discipline & FAQ).
  2. Follow 4–6 labs/authors: OpenAI, DeepMind, Anthropic, Meta AI, Stanford/Berkeley/ Tsinghua/Peking University, etc. (@author on X/GitHub).
  3. Look at model releases, not marketing: When a new model drops, read its technical report and evaluation card — not the social media headlines.
  4. Return to this page: Every quarter, cross-check against "nine major trends" for any new additions or reversals — updating both the paper map and this page is your personal frontier coordinate system.
  5. Check Model Compendium for new model profiles; dataset timeliness is subject to official releases.

About "frontier anxiety"

You don't need to read every new paper. The correct way to approach the frontier is "track only 1–2 milestones per tributary per week," and rely on surveys (surveys) and secondhand intermediaries for the rest. The criterion for judging whether a paper is worth deep-reading is in Reading Discipline & FAQ's "how to judge paper quality."

Further Reading ​

References ​

Key original papers and official resources referenced on this page (real links):