Appearance
The Research Frontier
"The frontier" is a trap word. On one side it means opportunity — between 2024 and 2025, nearly every quarter brought a paradigm-level shift; on the other, it means the densest information noise around — hundreds of new papers on arXiv every day, a new model release every week, and someone declaring "revolution" on every thread. Chase it without a method and you will burn out in the noise; ignore it entirely and you may miss the most important structural shift this field has seen in a decade.
This article focuses on the eight frontier threads most worth watching in 2024–2025 (data as of mid-2025), taking each apart in turn: where things stand, which papers and products represent them, and where they are stuck. It then delivers the article's most important judgment — what will last and what is hype — and closes with a low-cost, low-noise way to track the frontier. It carries forward the global perspective of the papers section hub and the paper map, and complements the classic paper deep-dives (which explain the "past") and reading discipline and FAQ (which explain "how to read").
A note on timeliness: model versions, leaderboards, and product progress move extremely fast. Facts in this article are current as of mid-2025 (dataAsOf 2025-06); for anything later, refer to the model and leaderboard quick reference.
Part 1: The Eight Frontier Threads (2024–2025)
Thread 1: Reasoning Models — "Reinforcement Learning + Chain-of-Thought" Becomes the New Paradigm
The state of play. In September 2024, OpenAI released o1, the first productization of "spending more compute at inference time (test-time compute)": the model uses reinforcement learning (RL) to learn to produce long internal chains of thought (CoT) — the longer it thinks, the more accurate the answer. o3 was previewed in December 2024, and o3/o4-mini shipped in April 2025; in January 2025, DeepSeek-R1 fully reproduced this route with open weights, proving that "RL teaching reasoning" is not the exclusive property of closed-source labs.
Unlike the earlier era of "prompts eliciting thinking," the core of this wave is the reinforcement learning signal during training: the model is first tuned with RL on domains where answers are verifiable, such as math and code, and behaviors like reflection, verification, and backtracking then emerge on their own. This technique and alignment (preference-learning methods like RLHF and DPO) are two sides of the same family — one teaches capability, the other teaches values. For a full mechanism breakdown, see the DeepSeek-R1 case study.
Representative work and products.
| Work / Product | Date | One-line takeaway |
|---|---|---|
| Chain-of-Thought Prompting (Wei et al.) | 2022-01 | Reasoning examples in the prompt make the model "think a step, answer a step," sharply improving math and commonsense reasoning |
| Let's Verify Step by Step (OpenAI) | 2023-05 | A process reward model (PRM) verifies reasoning step by step — the direct forerunner of reasoning models |
| o1 series (Learning to Reason with LLMs) | 2024-09 | Large-scale RL trains long chains of thought; competition math and code reach human-expert level |
| DeepSeek-R1 (DeepSeek) | 2025-01 | Pure RL with GRPO teaches reasoning; open weights, with the capability distilled into smaller models |
| Scaling LLM Test-Time Compute Optimally | 2024-08 | For verifiable tasks, "thinking longer at inference" can be more cost-effective than "making the model bigger" |
| o3 / o4-mini | 2024-12 / 2025-04 | Explicitly readable reasoning chains plus tool calling, linking "thinking" with "doing" |
Key open questions.
- The boundary of reward signals: RL needs verifiable answers (math, code). Writing, planning, and open-ended Q&A have no single right answer, and how to construct rewards there remains an open problem — the R1 paper itself admits this.
- The cost equation: "thinking longer" means inference costs rise by tens of times. Which tasks deserve deep thinking, and how to allocate thinking time adaptively, are the core engineering questions.
- Transparency of thought: Are long chains of thought really "reasoning," or just more elaborate pattern matching? For exploration at the mechanism level, see Thread 7 and safety and alignment.
One-line verdict: Reasoning models are the most certain structural trend of 2024–2025 — the "spend more inference compute to buy capability" curve is real on verifiable tasks; but "every task deserves deep thinking" is hype.
Thread 2: Long Context — Million-Token Windows and "Context Engineering"
The state of play. In February 2024, Google announced that Gemini 1.5 natively supported a 1-million-token context (technical report), and other labs have since stretched windows to 128K–2M. Once windows grew long, the question shifted from "does it fit?" to "can we afford to use it?": attention scales quadratically, the KV cache eats GPU memory, and positional encodings extrapolate poorly. Hence the surge of interest in sparse attention, KV cache compression, and positional encoding extrapolation. In early 2025, a DeepMind paper formally named "Context Engineering" — programming the context the way you write code, treating prompts, memory, and retrieval as tunable system resources. For the deployment and quantization details behind this, see inference optimization and quantization.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| Gemini 1.5 | 2024-02 | The first flagship model with million-token context — long documents, video, and codebases ingested in one go |
| Lost in the Middle (Liu et al.) | 2023-07 | Models "remember the beginning and end of long documents but drop the middle" — length ≠ effective use |
| YaRN | 2023-09 | Learned scaling of rotary positional encoding extends the RoPE window severalfold |
| LongRoPE | 2024-02 | Position interpolation stretches context to the million-token scale without fine-tuning |
| RingAttention | 2023-10 | Sequence-level parallelism that makes ultra-long-context training scalable |
| Infini-attention | 2024-04 | Compressive memory plus local attention, handling unbounded lengths in a single layer |
| Context Engineering for Neural Agents | 2025-01 | Named "context engineering": treating context as a programmable system resource |
Key open questions.
- Nominal vs. effective length: With a bigger window, can the model actually "use" what sits deep inside? "Lost in the Middle" shows that long-context utilization quality remains uneven.
- Cost and latency: The KV cache grows linearly with length; the cost and latency of million-token inference are real constraints.
- Evaluation gaps: Long-context benchmarks often test only "retrieving a single fact," leaving multi-hop reasoning and long-range consistency under-covered — for how to read benchmarks, see evaluation and benchmarks.
One-line verdict: Long context will last — it is the foundation for agents and multimodal memory; but the "indiscriminate infinite window" is hype — the endgame is a combination of sparsification, compression, and retrieval.
Thread 3: Efficient Architectures — State Space Models, Hybrids, and MoE
The state of play. In the decade of Transformer dominance, "attention is quadratic" has been a persistent pain point. At the end of 2023, Mamba used a state space model (SSM) to bring sequence-modeling complexity down to linear, sending shockwaves through the field. By 2024, the mainstream view was no longer "replace the Transformer" but "hybridize": Transformers handle precise global matching, while SSMs provide linear-cost long-range memory (Jamba, Mamba-2, and others). Meanwhile, mixture of experts (MoE) became the standard architecture for large models: a large total parameter count with only a few experts activated per token, trading sparse activation for capacity — Mixtral, DeepSeek-V3, and Llama 4 all follow this path. For the underlying mechanics, see the Transformer and attention.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| Mamba (Gu & Dao) | 2023-12 | Selective SSM: linear-complexity sequence modeling with clear advantages on ultra-long sequences |
| Mamba-2 | 2024-05 | Unifies SSMs under the attention framework — both fast and accurate |
| Jamba (AI21) | 2024-03 | The first production-grade Transformer + Mamba hybrid |
| Mixtral of Experts | 2024-01 | Sparse MoE: half the activated parameters still tops the leaderboards |
| DeepSeek-V3 | 2024-12 | A 671B-total / 37B-activated MoE trained at a fraction of the cost of comparable models |
Key open questions.
- The SSM ceiling: Pure SSMs still trail attention on ultra-long-range associations and precise retrieval, and there is no consensus on the right "mixing ratio."
- MoE's cost shift: MoE trades training cost for deployment cost — all experts must sit in GPU memory, and routing imbalance and batching efficiency remain active research areas.
- The sustainability of sparsity: Does activating fewer parameters lower the model's capability ceiling? Still unsettled.
One-line verdict: Hybrid architectures and MoE will last — they are the mainstream solution for scaling models under finite compute; "SSMs fully replacing attention" is hype, and over the next five years we are more likely to see the two divide the labor within layered architectures.
Thread 4: Unified Multimodality — "Any Input, Any Output"
The state of play. Multimodality has moved from dual-tower "vision + language" alignment (CLIP) to generative vision-language models (LLaVA, GPT-4V), and in 2024 it took another step: natively multimodal training — treating text, images, audio, and video as a single token stream from the very first line of code, rather than bolting on a vision encoder afterward. The representatives are OpenAI's GPT-4o (2024-05) and Google's Gemini series (natively multimodal from 2.0 on). The 2025 direction is any-modality input → any-modality output ("any-to-any"): one model handles text-to-image, image-to-text, and audio/video understanding and generation. For principles and trade-offs, see multimodal models.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| CLIP | 2021-01 | Contrastive image-text learning unified the semantic space — the infrastructure of multimodality |
| Flamingo / BLIP-2 | 2022-04 / 2023-01 | Frozen vision encoders + lightweight connector layers deliver visual capability at low cost |
| Visual Instruction Tuning (LLaVA) | 2023-04 | Instruction tuning aligns visual features into the language model — the baseline for open-source VLMs |
| GPT-4o | 2024-05 | End-to-end native multimodality; voice conversation latency drops to human levels |
| Gemini 2.0 | 2024-12 | Native multimodality plus native tool use and image generation, on the road to agents |
Key open questions.
- Depth of alignment: VLM "understanding" often stops at semantic alignment (being able to say "what this is"), still far from pixel-level understanding (precise spatial relations, counting, physical intuition).
- Evaluation and methodology: Multimodal capability boundaries are blurry, and benchmarks for hallucination rates and instruction following are still evolving rapidly.
- Data and compute: There is no standard answer for any-to-any data mixtures or tokenization schemes, and training costs multiply.
One-line verdict: Native multimodality will last — it is absorbing the three separate tracks of speech, vision, and language; "more modalities = more intelligence" is hype — more modalities do not automatically mean stronger reasoning, and may amplify hallucinations instead.
Thread 5: Agents and Tools — MCP, Computer Use, and Multi-Agent Systems
The state of play. From late 2024 into 2025, agents moved from "demos" to "engineering standards": in November 2024, Anthropic open-sourced the Model Context Protocol (MCP), defining a unified interface between models ↔ tools/data; from October 2024, Anthropic and OpenAI successively launched Computer Use — letting models operate browsers and desktops the way humans do; and in March 2025, general-purpose agent products such as Manus went viral across the Chinese internet. Multi-agent collaboration became a research hotspot: splitting one task among several role-specialized agents, equipped with shared memory and communication protocols. For a systematic understanding, see AI agents; for product cases, see Manus and agent applications.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| ReAct | 2022-10 | Alternating reasoning and action — the archetypal agent thought pattern |
| Toolformer | 2023-02 | Models teach themselves to call tools; tool use goes mainstream |
| AgentBench | 2023-08 | Multi-environment agent evaluation, charting the gap between LLMs and humans |
| SWE-bench | 2023-10 | A benchmark of real GitHub issues that makes "agents writing code" measurable |
| Model Context Protocol (MCP) | 2024-11 | A unified protocol for tool/data access — the "USB-C port" for agents |
| Anthropic Computer Use / OpenAI CUA (Operator) | 2024-10 / 2025-01 | Models operating screens and software directly — from "chatting" to "doing" |
| Building Effective Agents (Anthropic) | 2024-12 | Official engineering wisdom: if simple works, don't reach for complex multi-agent setups |
Key open questions.
- Reliability is the survival line: For a 10-step task, 95% per-step success means only about 60% overall. There is no silver bullet for the stability of long-horizon tasks (dozens of steps plus memory).
- Evaluation lag: Benchmarks like SWE-bench saturate quickly, and evaluation of open-ended real-world tasks remains a blank spot.
- Multi-agent vs. single agent: Is dividing roles among several agents better than one strong agent? Experimental results in 2025 increasingly favor "most of the time, a single agent with good tools is enough."
One-line verdict: Tool calling, MCP, and Computer Use will last — they are the standard infrastructure of agents; "complex multi-agent frameworks" are most likely hype — industry consensus is returning to "simple code + a smart model."
Thread 6: Video and World Models — Sora and "Predicting Physics"
The state of play. In February 2024, OpenAI released the Sora technical report, Video generation models as world simulators, demonstrating long-shot consistency and rudimentary physical intuition; in December, Sora 2 launched officially, with major gains in duration, controllability, and consistency. Around the same time, DeepMind's Genie generated interactive worlds from a single image, turning "world model" from a sci-fi term into an engineering direction: using generative models to learn the world's spatiotemporal regularities and thereby predict what comes next. The combination of diffusion models and Transformers (DiT) is the general-purpose foundation underneath. For principles, see diffusion models and generative AI; for the full case, see Sora and video generation.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| Scalable Diffusion Models with Transformers (DiT) | 2022-12 | The Transformer-ized diffusion model — the general foundation for text-to-image/video |
| Sora (technical report) | 2024-02 | Proclaimed a "world simulator"; a milestone in long-shot consistency |
| Genie (DeepMind) | 2024-02 | Interactive 2D worlds from a single image — an interactive world model |
| Sora 2 | 2024-12 | Controllable generation moving from "single shots" to "multi-shot + temporal consistency" |
| JEPA / the world-model route (LeCun) | 2022– | Self-supervised prediction of representations — a different world-model path from generative reconstruction |
Key open questions.
- Physical consistency: Object permanence and causal interactions still break down — the debate between "emergent physics" and "data interpolation" remains unresolved.
- Two world-model routes: Whether generative (diffusion reconstructing pixels) or predictive (JEPA predicting representations) is closer to "understanding the world" is still unsettled.
- Cost and controllability: The inference cost of video generation and frame-level controllability remain barriers to deployment.
One-line verdict: Video generation will last — it is the reset switch for the content industry; "world models already understand physics" is hype — for now they are better described as "very good at interpolating video data," still far from causal world models.
Thread 7: Safety and Alignment Research — Interpretability and Automated Red-Teaming
The state of play. The stronger models get, the more "how do we make them reliable, controllable, and interpretable" becomes a mainline concern. 2024–2025 brought two notable advances: mechanistic interpretability moved from "visualizing features" to "reading the model at scale" — Anthropic's dictionary learning scaled from a small 4-layer model to tens of millions of features inside flagship models, and then in 2025 cross-layer features (Crosscoders) provided a first localization of the "neural coordinates of concepts"; and automated red-teaming — models attacking each other and automatically hunting for vulnerabilities, replacing most manual red-teaming. Alignment methods themselves have also branched out from RLHF into DPO, RLAIF, Constitutional AI, and more. For the full picture, see AI safety and governance.
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| InstructGPT (RLHF) | 2022-03 | Reinforcement learning from human feedback — the paradigm starting point of alignment |
| Constitutional AI | 2022-12 | AI self-critique + principle-based constraints; automation of alignment and red-teaming |
| Red Teaming Language Models | 2022-09 | Automated red-teaming: using models to find flaws in models |
| Direct Preference Optimization (DPO) | 2023-05 | Preference optimization without a reward model — a major simplification of training |
| Towards Monosemanticity (dictionary learning) | 2023-10 | Sparse autoencoders decompose a model's internals, surfacing "single-meaning" features |
| Scaling Monosemanticity (Golden Features) | 2024-05 | Pushing mechanistic interpretation to the scale of flagship-model internals |
| Mapping the Mind of a LLM (Crosscoders) | 2025-02 | Cross-layer features: locating "where in the network" a concept lives |
| Humanity's Last Exam | 2025-01 | A high-difficulty capability benchmark: a splash of cold water on the "models have surpassed experts" narrative |
Key open questions.
- Interpretability's scale bottleneck: Plenty of features have been found, but "understanding how they combine into behavior" remains handcraft work.
- Provable safety: Every defense to date is empirical, with no formal guarantees; red-teaming can prove "a vulnerability was found" but never "there are none."
- The tension between technical alignment and governance: How training methods (technical alignment) and regulatory audits (governance) work together is the most concrete question of 2025.
One-line verdict: Interpretability and automated red-teaming will last — they are irreversible infrastructure; "interpretability can already explain large models" is hype — we are still in the "reading features" stage, far from "reading minds."
Thread 8: From "Conversation" to "Autonomous Output" — Agentic RL, Vibe Coding, and the AI Scientist
The state of play. The clearest signal of 2025: AI no longer just "answers" — it has started "doing work," and the way it works is itself being trained. Three directions converge here. Agentic RL upgrades reinforcement learning from "picking an answer" to "executing multi-step tasks" — DeepSeek-R1 used large-scale RL to learn "think it through, then answer" (see DeepSeek-R1 and reasoning models); the late-2024 Agentic RL paper (arXiv:2501.03266) went further, using RL to train an agent's dual abilities of search and reasoning, and in 2025 multiple institutions followed suit by training planning and tool calling with RL. Vibe Coding, coined by Karpathy in February 2025, describes a programming style of "describing intent in natural language, letting AI write the code, with humans only reviewing"; paired with agent-style tools such as Claude Code and Cursor, code production is shifting from "hand-writing" to "directing + reviewing" (see GitHub Copilot and code intelligence). And the AI Scientist automates the entire research pipeline of "read literature → form hypotheses → run experiments → write papers" (Sakana AI's AI Scientist-v2, August 2024), opening the "research agent" narrative as a complement to AlphaFold-style single-point breakthroughs (see AlphaFold and AI for Science).
Representative work.
| Work | Date | One-line takeaway |
|---|---|---|
| Agentic RL (arXiv:2501.03266) | 2024-11 | Systematically training agents with both "search + reasoning" abilities via RL — the starting point of the agentic RL paradigm |
| DeepSeek-R1 | 2025-01 | The milestone of large-scale RL for reasoning models; its open-source release ignited the reasoning-model wave |
| The AI Scientist-v2 | 2024-08 | An end-to-end automated research agent that independently closes the loop of "literature → experiments → paper" |
| Claude Code / Cursor | 2024–2025 | The industry representatives of agent-style coding tools — where Vibe Coding lands |
Key open questions.
- Reward design: How to assign rewards on sparse, long-horizon agentic tasks, and how to prevent "shortcut-taking" (reward hacking), remain open problems.
- Code quality and security: Vibe Coding removes the barrier to "writing code," but the burden of review grows heavier — code hallucinations, supply-chain risks, and privilege-escalation risks all need humans as the last line of defense (see common pitfalls and anti-patterns).
- Research ethics: The "paper mills" produced by AI Scientists raise controversies over reproducibility, authorship, and research integrity.
One-line verdict: Agentic RL will last — it unifies "reasoning" and "action" into a single training objective, the key to making agents reliable; "Vibe Coding replaces programmers" is hype — what it changes is how code gets written, not the need for people who understand code; the AI Scientist will last but pay off slowly — the bottleneck of research is not "writing papers" but "designing experiments and judging what matters."
Part 2: What Will Last and What Is Hype
All eight threads compressed into a single judgment table — the part of this article most worth taking away:
| Thread | Likely to last | Likely hype | One-line verdict |
|---|---|---|---|
| Reasoning models | RL + CoT, test-time compute, process rewards | "Every task deserves deep thinking" | A certain structural trend, realized only on verifiable tasks |
| Long context | Sparse attention, KV compression, retrieval, context engineering | "Infinite window + dump everything in at once" | Will last, but as an engineering combination rather than brute force |
| Efficient architectures | MoE, hybrid architectures (Transformer + SSM), distillation | "SSMs completely replace attention" | The inevitable choice under finite compute; the replacement narrative doesn't hold |
| Unified multimodality | Native multimodality, any-to-any, unified token streams | "More modalities means more intelligence" | Will absorb the specialized tracks, but more modalities don't automatically mean more capability |
| Agents and tools | MCP, Computer Use, tool calling, agentic coding | Complex multi-agent frameworks | The infrastructure is certain; complex frameworks are receding |
| Video and world models | Video generation, controllable generation, diffusion models | "World models already understand physics" | The content revolution is certain; physical understanding remains undelivered |
| Safety and alignment | Mechanistic interpretability, automated red-teaming, alignment algorithms | "Interpretability can already explain models" | Irreversible infrastructure; progress is overestimated |
| Agentic RL and the AI Scientist | The agentic RL training paradigm, agent-style coding tools | "Vibe Coding replaces programmers," "fully automated AI research" | The training paradigm will last; product narratives will pay off slowly |
A One-Line Rule of Thumb
To judge what will last, check whether it serves as the foundation for capabilities others can't do without: MCP is a foundation, multi-agent frameworks are not; sparse attention is a foundation, infinite context is not; RL teaching reasoning is a foundation, "just think a bit longer on everything" is not. Foundations stay; the narratives attached to them recede.
Part 3: How to Track the Frontier
The problem with frontier information is not scarcity but surplus. What works is not "read everything" but building a prioritized signal pipeline.
1. arXiv: The Primary Battlefield — Subscribe to Authors, Not the Listings
| Category | Topic | Notes |
|---|---|---|
| cs.LG | Machine Learning | The main battlefield, with the broadest coverage |
| cs.CL | Computation and Language / NLP | The home turf of LLM papers |
| cs.CV | Computer Vision | Where multimodal and video-generation papers often appear |
| cs.AI | Artificial Intelligence | Agents, reasoning, and general methods |
Hundreds of papers arrive every day, so scrolling the raw listings is pointless. Following authors beats following listings: subscribe on arXiv to new submissions from the authors of papers you have closely read; use Hugging Face Papers to aggregate by popularity, and arXiv Sanity to filter by similarity. Close-read only high-signal papers: the title speaks to your own questions, the author has a consistent track record, and code and weights ship alongside.
2. X/Twitter: Researcher Activity Is the Earliest Signal
Frontline researchers often stake out positions before the paper is even written — worth following over the long term:
- Andrej Karpathy (@karpathy): architectures, training details, deep learning education;
- Yann LeCun (@ylecun): the world-model route, JEPA;
- Sebastian Raschka (@rasbt): paper breakdowns and implementation details;
- Chip Huyen (@chipro) and Simon Willison (@simonw): engineering and product perspectives;
- Official lab accounts (OpenAI, Anthropic, Google DeepMind, Meta AI, DeepSeek, Hugging Face): first-party announcements.
Signal vs. Noise
Virality on X correlates poorly with importance: marketing, emotional debates, and clickbait make up most of the volume. Keep your follow list under 20, and treat what they say as "hypotheses to verify" rather than conclusions — what truly shapes your judgment is the papers you have read and the experiments you have run yourself.
3. Top Conferences: Systematized, Distilled Knowledge
| Conference | Field | Roughly each year |
|---|---|---|
| NeurIPS | Machine learning (general) | December |
| ICML | Machine learning | July |
| ICLR | Representation learning / deep learning | April–May |
| ACL / EMNLP | Natural language processing | July / November |
| CVPR | Computer vision | June |
No need to attend: skim the titles when the accepted-papers list drops, and once the conference ends, go straight to tutorials and surveys — one or two high-quality tutorials beat fifty papers, because they lay out the coordinate system of the field. For the complete reading method, see the paper reading path.
4. Evaluation Sites: Trading "Feelings" for "Numbers"
| Channel | What it's for |
|---|---|
| LMArena | Blind human-vote head-to-head battles; among the most trustworthy leaderboards for overall capability |
| Papers with Code | Task → SOTA → paper → code |
| Artificial Analysis | Puts "how capable" side by side with "how costly and how fast" |
| Hugging Face Papers | Daily paper trending + code + discussion |
Fast signal ┌──────────────────────────────────────────────┐
│ X/blogs, lab news ── signals, noisiest │
│ Daily arXiv updates ── primary papers │
│ Leaderboards / HF Papers ── aggregation + numbers │
│ Conference papers, tutorials ── distilled knowledge │
Deep signal └──────────────────────────────────────────────┘
The deeper you read, the steadier your judgment;
the more urgently you chase, the easier you get misled.A Suggested Weekly Rhythm
Twenty minutes a day: in the morning, scan HF Papers titles and reshares from the researchers you follow, and close-read just one paper. Half a fixed day each month: read one survey and check LMArena and the model leaderboards. Spend the remaining time on reproducing results and writing notes — that is where judgment truly accumulates. For the complete long-term plan, see the learning paths overview.
Further Reading
- Start Here — the main entry to the papers section, with search and reading methods
- The Paper Map — grasp the coordinates from classics to frontier in one chart
- Classic Paper Deep-Dives — close readings of the papers cited here that have become classics
- Reading Discipline and FAQ — the three-pass reading method, the red-flag checklist, and common questions
- The DeepSeek-R1 Case Study — a full breakdown of the reasoning-model paradigm
- Sora and Video Generation — the full story behind the world-model narrative
- The Transformer and Attention — where every architectural evolution begins
- Multimodal Models — the mechanics and trade-offs of unifying modalities
- AI Agents — the principles of tool calling and multi-agent systems
- AI Safety and Governance — alignment, interpretability, and red-teaming
- Inference Optimization and Quantization — making long context and inference costs practical
- LLM Evaluation and Benchmarks — how to read and use benchmarks
- A Brief History of the Field — where the frontier sits in the timeline of past years
- Model and Leaderboard Quick Reference — tracking model versions and leaderboards
- Curated Resource List — a toolbox for tracking the frontier
References
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv:2501.12948)
- OpenAI. Learning to Reason with LLMs (the o1 series)
- OpenAI. Introducing o3 and o4-mini
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (arXiv:2201.11903)
- Let's Verify Step by Step (arXiv:2305.20050)
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (arXiv:2408.03314)
- Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context (arXiv:2403.05530)
- Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172)
- YaRN: Efficient Context Window Extension of Large Language Models (arXiv:2309.00071)
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens (arXiv:2402.13753)
- RingAttention with Blockwise Transformers for Near-Infinite Context (arXiv:2310.06236)
- Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention (arXiv:2404.07143)
- Context Engineering for Neural Agents (arXiv:2501.17437)
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality (Mamba-2, arXiv:2405.21060)
- Jamba: A Hybrid Transformer-Mamba Language Model (arXiv:2403.19887)
- Mixtral of Experts (arXiv:2401.04088)
- DeepSeek-V3 Technical Report (arXiv:2412.19437)
- OpenAI. Hello GPT-4o
- Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv:2103.00020)
- Visual Instruction Tuning (LLaVA, arXiv:2304.08485)
- OpenAI. Video Generation Models as World Simulators (Sora technical report)
- OpenAI. Sora 2
- Genie: Generative Interactive Environments (arXiv:2402.15391)
- Scalable Diffusion Models with Transformers (DiT, arXiv:2212.09748)
- Model Context Protocol documentation
- Anthropic. Introducing Computer Use (Claude 3.5)
- OpenAI. Introducing the Computer-Using Agent (Operator)
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv:2310.06770)
- AgentBench: Evaluating LLMs as Agents (arXiv:2308.03688)
- Anthropic. Building Effective Agents
- Anthropic. Towards Monosemanticity (Transformer Circuits)
- Anthropic. Scaling Monosemanticity (Golden Features, Transformer Circuits)
- Mapping the Mind of a Large Language Model (Crosscoders, arXiv:2502.16021)
- Training Language Models to Follow Instructions with Human Feedback (InstructGPT, arXiv:2203.02155)
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (arXiv:2305.18290)
- Constitutional AI: Harmlessness from AI Feedback (arXiv:2212.08073)
- Red Teaming Language Models to Reduce Harms (arXiv:2209.07858)
- Humanity's Last Exam (arXiv:2501.14249)
- arXiv (latest cs.LG / cs.CL submissions)
- Hugging Face Papers
- Papers with Code
- LMArena
- Artificial Analysis
- arXiv Sanity Preserver