Appearance
Frontier Advances
"Frontier" is a trap word. It means opportunity, but also the densest information noise — dozens of inference papers flood arXiv every week, a new engine ships every month, and every research line has someone declaring a "3x speedup." Chase it without method and you will exhaust yourself; ignore it completely and you will miss the field's once-in-a-decade structural shifts (FP8 in the H100 era, FP4 in the B200 era, prefill/decode disaggregation, the training-time internalization of speculative decoding).
The goal of this page is a map with method: first, how to track the frontier at low cost and low noise (channels and priorities); then, six research lines most worth watching today (status, representative work, key questions); then, an honest answer to "which problems are still unsolved"; and finally, concrete advice on choosing a direction to go deep. It expands on Start Here and complements Classic Papers in Depth and Reading Discipline & FAQ.
1. How to Track the Frontier
Frontier information is not too scarce — it is too abundant. What effective trackers do is not "read everything" but build a prioritized signal pipeline.
1.1 arXiv: The First-Hand Battlefield
Almost every important work appears on arXiv as a preprint before it is accepted anywhere. For inference-acceleration researchers, a few subcategories are the main battlefield:
| Category | Topic | Notes |
|---|---|---|
| cs.LG | Machine learning | Operator, quantization, and speculative-decoding papers often appear here |
| cs.DC | Distributed computing | Serving-system, parallelism, and cluster-scheduling papers often appear here |
| cs.OS | Operating systems | SOSP/OSDI submissions often appear here first |
| cs.PF | Performance | Benchmark, roofline, and profiling work |
| cs.AR | Computer architecture | Hardware co-design and ISA papers |
Daily volume is in the dozens of papers, so scrolling the raw list is extremely inefficient. Three practical habits:
- Follow authors and institutions: on arXiv, subscribe to new submissions from scholars you follow (for example, authors of papers you have closely read) rather than scanning the full list. A productive group usually beats a hundred random papers. Key authors: Tri Dao (Princeton/Together), Lianmin Zheng (Berkeley/DeepSeek), Ying Sheng (MIT/Together), Zhuoming Chen (Princeton), the DeepSeek-AI systems group, NVIDIA's TensorRT-LLM team, the vLLM core team.
- Use tools to filter: Hugging Face Papers curates hot papers daily with code and discussion; Papers with Code finds SOTA by task; tools like AlphaXiv let you ask questions directly on a paper's page.
- Only deep-read high-signal papers: to judge whether a paper deserves a deep read, check three things first — does the title answer a deployment problem you are working on, do the authors/institutions have a track record, and did they release code and weights.
A timeline lesson
Preprints always beat conferences: submission to publication is often 6-10 months apart. The vLLM paper was posted on arXiv in June 2023 and only appeared at SOSP in October 2023, but its influence began the day the preprint dropped. Read preprints and verify against the conference version — the standard rhythm of researchers.
1.2 Benchmarks and Leaderboards: Turning "Feelings" into Numbers
Chasing papers alone lets narratives lead you astray — press releases always report good news. The second pipeline against noise is benchmark aggregators:
| Channel | Use | Characteristics |
|---|---|---|
| Papers with Code | Find SOTA leaderboards by task | Highest score, code, and paper mapped one-to-one for every benchmark |
| Artificial Analysis | Inference cost / speed / quality comparison | Puts "how strong" and "how expensive" side by side, close to engineering decisions |
| vLLM benchmarks | Engine-level throughput / latency benchmarks | Evolves with vLLM main, closest to industrial deployment |
| SGLang benchmarks | Engine comparison + speculative-decoding benchmarks | Comparison data for speculative decoding and RadixAttention |
Know the limits of benchmark sites too: leaderboards always lag releases and are vulnerable to score manipulation. Their correct use is not "whoever is first is trustworthy" but giving you a horizontal coordinate to judge whether a claimed improvement in a paper is real or noise — the method and pitfalls are in the reading method of Start Here, and in Curated Resources.
1.3 Open-Source Repos and PRs: The Second Pipeline of Engineering Signals
Inference deployment has a distinctive phenomenon: much important work appears in GitHub PRs before it becomes a paper. Repos worth watching long-term:
| Repo | Value |
|---|---|
| vllm-project/vllm | Mainline PRs hide the evolution of PagedAttention, the landing of chunked prefill, and the integration of various speculative-decoding schemes |
| NVIDIA/TensorRT-LLM | First-hand signal of engineering-hardware co-design; FP8/FP4 kernels usually land here first |
| flash-attention | The official implementation of the FA family and related operators; issues contain many edge cases |
| flashinfer-ai/flashinfer | The official implementation of composable KV formats and block-sparse attention |
| sgl-project/sglang | The engineering frontier of RadixAttention and speculative decoding |
| llama.cpp | The on-device inference de facto standard, and the evolution of the GGUF format |
1.4 Top Conferences: Settled Systems Knowledge
Conference papers arrive later than preprints but have survived peer review, and conferences organize tutorials and surveys — the best entry points for systems knowledge:
| Conference | Field | Rough timing each year |
|---|---|---|
| SOSP / OSDI | Operating systems and systems | Alternate Oct-Nov |
| MLSys | Machine learning systems | March |
| ASPLOS | Hardware-software interface | March-April |
| ISCA / HPCA | Architecture | June / April |
| NeurIPS / ICML / ICLR | Machine learning general | December / July / April-May |
You do not need to attend to track conferences: when the acceptance list drops, scan the titles; after the conference, read the tutorial materials. For a newcomer, one or two high-quality tutorials beat fifty papers, because they lay out the field's coordinate system.
text
Arrange the four pipelines by "speed x depth":
High speed ┌────────────────────────────────────────────────────────────┐
│ GitHub PRs, X/blogs, lab updates -- signal, noisiest │
│ Daily arXiv updates -- first-hand papers │
│ Leaderboards / HF paper daily -- aggregation + numbers │
│ Conference papers & tutorials -- settled knowledge │
High depth └────────────────────────────────────────────────────────────┘
The deeper you read, the steadier your judgment; the faster you chase, the easier you are misled.2. Six Research Lines
The six lines below are the most active and most industry-shaping directions today. Each gives status, representative work, and key questions, and cites work as "paper name (arXiv number) -> one-line takeaway" so you can cross-read them in Classic Papers in Depth.
Line 1: Operator Frontier — From IO-Awareness to Hardware Co-Design
Status. FlashAttention v1/v2 corrected the attention bottleneck from FLOPs back to IO; since 2024 the operator frontier has entered its third generation: deep co-design with hardware generations. FlashAttention-3 uses H100's FP8 + asynchrony (warp-specialized) to push the peak to 1.2 PFLOPs; FlashInfer uses block sparsity + composable KV formats to unify paged KV and variable-length attention in one kernel; vAttention asks the reverse question: can hardware-native sparse attention replace software PagedAttention and avoid SM stalls?
Representative work.
| Work | Takeaway |
|---|---|
| FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (arXiv:2407.08608) | FP8 + asynchrony (warp-specialized) on H100, 1.2 PFLOPs peak, 1.5-2x faster than FA2; shows the paradigm of "operators co-designed with a hardware generation" |
| FlashInfer: An Efficient and Customizable Attention Engine for LLMs (arXiv:2401.12448) | Block-sparse + composable KV formats (PagedKV, RaggedKV, BlockSparse) unifying the attention interface; a 2025 MLSys paper |
| vAttention: Efficient Attention with Suffix Parallelism (arXiv:2405.04447) | Hardware-native sparse attention replaces software PagedAttention, avoiding SM stalls; argues "software paging is a stopgap, hardware sparsity is the endgame" |
| Flash-Decoding (Dao et al., 2023 blog) | For long-sequence, batch=1 decode, an FA parallelism improvement that splits the KV into parallel chunks |
| Lightning Attention-2 (Qwen team, 2024) | An engineering implementation of linear attention, with inference cost near-linear in long contexts |
Key questions.
- The hardware generation decides the algorithm choice: FA2 is optimal on A100, FA3 on H100, and FP8/FP4 kernels are only sweet on Blackwell. Change the hardware generation and the operator must be rewritten — the core pain point of the operator layer.
- Software vs. hardware sparsity: vAttention's "replace software paging with hardware sparsity" is the most interesting debate of 2024. Software schemes remain mainstream short-term, but whether hardware sparsity will end PagedAttention is an open question worth tracking.
- Missing unified interface: FlashInfer's attempt shows there are too many attention-operator interfaces (PagedKV, RaggedKV, BlockSparse, FA) and no unified abstraction. A daily pain point of engine developers. Operator mechanics: Kernel Fusion and Custom Kernels and GPU Architecture and Optimization.
How this relates to you
If you build operators: the operator layer is a "must-track" line — a hardware-generation shift means rewriting the algorithm. If you build engines: watch only interfaces and benchmarks, not every paper. If you do research: the operator layer is the hardest but highest-impact direction — an FA4-level work is a best-paper candidate at NeurIPS/MLSys.
Line 2: Serving-Systems Frontier — From Single Engines to Prefill/Decode Disaggregation
Status. vLLM's PagedAttention solved KV cache fragmentation, but LLM serving has two leftover problems: (1) prefill and decode have exactly opposite compute/memory ratios and interfere with each other on the same GPU in the same batch; (2) in multi-turn dialogue and few-shot scenarios, prefix recomputation of the KV cache wastes work. The two main lines of 2024 — physical prefill/decode disaggregation and prefix-tree reuse — address these two problems respectively.
Representative work.
| Work | Takeaway |
|---|---|
| DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (arXiv:2401.09670) | Physically separate prefill and decode onto different GPUs, each scaling independently; P99 latency halved |
| Splitwise: Efficient generative LLM inference using phase splitting (HPCA 2024) | Similar to DistServe but focused on cluster-level scheduling of phase-specific machines |
| SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills (arXiv:2308.16369) | Chunked prefill: split a long prefill into chunks sharing a batch with decode, balancing compute and memory; integrated in vLLM 0.5+ |
| SGLang: Efficient Execution of Structured Language Model Programs (arXiv:2312.07104) | RadixAttention: prefix-tree reuse of the KV cache, a throughput weapon for multi-turn dialogue and few-shot scenarios |
| Mooncake: A KV-Centric Decoupled Architecture for LLM Serving (arXiv:2407.00079) | Moonshot AI's planet-scale architecture: stores and transfers the KV cache as a first-class citizen |
| AlpaServe: Scaling for Large-Scale Deep Learning Serving (OSDI 2023) | Statistical multiplexing analysis for multi-model deployment, guiding time-multiplexing |
Key questions.
- Disaggregation vs. mixing: DistServe's physical disaggregation costs uneven machine utilization (prefill machines busy, decode machines idle); SARATHI's mixing costs prefill/decode interference. Short-term, mixed schemes (SARATHI) ship more; long-term, disaggregation (DistServe) is more likely to become the paradigm for large-scale serving.
- KV cache storage and transfer: Mooncake treats the KV cache as a first-class citizen, but the bandwidth cost and latency overhead of cross-node KV transfer remain open. Will "distributed storage of the KV cache" become the next vLLM-level paradigm? Worth sustained attention.
- Prefix reuse vs. chunked prefill: SGLang's RadixAttention and SARATHI's chunked prefill both help and conflict in multi-turn dialogue. Mainstream engines are merging the two, but best practices are still evolving. Serving mechanics: Batching and Request Scheduling and Model Serving and Orchestration.
Line 3: Speculative-Decoding Frontier — From 50% Acceptance to 6.5x Speedup
Status. When Leviathan formalized speculative decoding in 2023, the acceptance rate was ~50% and the speedup 2-3x. In 2024 Medusa/EAGLE-1/2 pushed the acceptance rate to 70%+ and the speedup to 3-5x. In 2025 EAGLE-3 pushed the speedup to 6.5x. Meanwhile, DeepSeek-V3 internalized speculative decoding into the training objective (MTP), opening the new paradigm of "learning to draft at training time." Speculative decoding is one of the most active research directions of LLM inference in 2024-2026.
Representative work.
| Work | Takeaway |
|---|---|
| EAGLE-3: Scaling up Inference Speed (arXiv:2503.01840) | NeurIPS 2025; deeper features + a training-objective improvement, 3-6.5x speedup; the latest benchmark of the EAGLE series |
| P-EAGLE (AWS, 2025) | Decoupled drafting: the draft model can be deployed independently, decoupled from the main model, simplifying multi-model integration |
| SpecForge (SGLang team, 2025 blog) | A multi-draft-model training pipeline: treating speculative decoding as an "engineering product" that drafts can be mass-produced |
| DeepSeek-V3 MTP (arXiv:2412.19437) | The main model natively supports multi-token prediction, learning speculative-decoding ability at training time — the training-time internalization of speculative decoding |
| Medusa (arXiv:2401.10774) | Main-model multi-head parallel prediction, simplifying speculative decoding; the prelude to EAGLE |
| EAGLE-2 (arXiv:2406.16858) | Dynamic draft tree, 70%+ acceptance rate |
| SpecBench (SGLang, 2024) | A unified evaluation benchmark for speculative decoding, comparing EAGLE/Medusa/SpecInfer and other schemes |
Key questions.
- Training-time vs. inference-time drafting: DeepSeek-V3 MTP's "learn to draft at training time" vs. the EAGLE series' "attach a draft at inference time" — which becomes mainstream? Short-term EAGLE is more productionized; long-term MTP, deeply coupled with model architecture, may win. The key open question of 2025-2026.
- Draft-model decoupling and portability: P-EAGLE and SpecForge make the draft model an "independently trained artifact," meaning there may be "draft-model-as-a-service" in the future — a draft model shipped alongside a new model release may become standard.
- The ceiling of the acceptance rate: EAGLE-3 reaches 80%+ acceptance and 6.5x speedup. Can the acceptance rate reach 100%? Obviously not — but where is the ceiling, and can an adaptive budget (guess more on easy tokens, less on hard ones) work, is an active research direction. Cases: Speculative Decoding and Medusa/EAGLE.
Line 4: Quantization Frontier — From 4-bit to FP4 and Rotation Quantization
Status. The quantization sweet spot moved from INT4 (GPTQ, AWQ) → INT8 activations (SmoothQuant) → FP8 in 2024 (H100) → FP4 in 2025 (Blackwell B200). Meanwhile, rotation quantization (QuaRot, SpinQuant) uses equivalent transforms to "flatten" distributions, greatly improving W4A4 accuracy. The two main lines of quantization — hardware formats and equivalent transforms — are converging.
Representative work.
| Work | Takeaway |
|---|---|
| FP8 Formats for Deep Learning (NVIDIA, white paper) | H100 introduces FP8 Tensor Cores, a hardware-level sweet-spot format; E4M3 and E5M2 |
| FP4 Inference on Blackwell (NVIDIA, 2024) | B200 introduces FP4 Tensor Cores + the NVFP4 format, a weight-quantization sweet spot |
| QuaRot: Outlier-free 4-bit Inference for Rotated LLMs (arXiv:2404.00456) | Random orthogonal transforms "flatten" outliers, greatly improving W4A4 accuracy |
| SpinQuant: LLM Quantization with Learned Rotations (arXiv:2405.16406) | Learned rotation matrices, better than QuaRot's random rotation, W4A4 approaching W4A16 |
| BigVGQ: Bit-Grouped Quantization (2024) | Quantize weights in subgroups, further reducing quantization error |
| GPTQ/AWQ/SmoothQuant (close reading here) | The classic 4-bit and 8-bit schemes based on second-order/activation/smoothing |
Key questions.
- Hardware formats vs. algorithmic quantization racing: FP8 makes 8-bit floating point hardware-native, but algorithmic quantization (GPTQ/AWQ) reaches 4-bit. Once hardware natively supports FP4/NVFP4 (B200), does algorithmic quantization still have an edge? Short-term GPTQ/AWQ remain mainstream on A100/H100; long-term NVFP4 may end algorithmic research on weight quantization.
- W4A4 vs. W4A16 accuracy gap: W4A16 (4-bit weights, 16-bit activations) is near-lossless, but W4A4 (all 4-bit) still loses noticeably. Rotation quantization (QuaRot, SpinQuant) brings W4A4 close to W4A16, but not quite. When will W4A4 become the default? Depends on a double breakthrough of hardware and algorithms.
- The theoretical limit of rotation quantization: QuaRot uses random rotation, SpinQuant learned rotation — what is the optimal rotation? Still open. In theory the "Hadamard transform" may be better; engineering is still validating. Quantization mechanics: Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
Line 5: Architecture Innovation — SSM, MoE, and Linear Attention
Status. The Transformer is not the only choice. Mamba/SSM has near-linear inference cost on long sequences; MoE (DeepSeek-V3, Mixtral) pushes open-source models to a hundred billion parameters; Linear Attention (Lightning Attention, GLA) tries to break the O(n²) complexity. Architecture-layer innovation is reshaping inference-engine design — MoE requires expert parallelism and routing, Mamba requires stateful inference, and Linear Attention requires chunk-wise computation.
Representative work.
| Work | Takeaway |
|---|---|
| Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752) | Selective state-space model, near-linear inference cost on long sequences, challenging the Transformer's O(n²) |
| Mixtral of Experts (arXiv:2401.04088) | Open-weights 8-expert MoE, bringing MoE from closed to open source; expert parallelism becomes standard on the inference side |
| DeepSeek-V3 Technical Report (arXiv:2412.19437) | 256 experts + 8 active MoE + MTP, the open-source benchmark of MoE inference serving |
| Lightning Attention-2 (Qwen team, arXiv:2401.11461) | An engineering implementation of linear attention, near-linear inference cost in long contexts |
| Gated Linear Attention (GLA, Yang et al., 2024) | Gated linear attention, easier to parallelize in training than Mamba |
| Jamba (AI21 Labs, 2024) | Mamba + Transformer hybrid architecture, the first industrial large-scale LLM to use SSM |
Key questions.
- The engineering challenge of MoE inference: memory residency of 256 experts, expert-routing load balancing, and cross-node communication are far more complex than Llama-3-70B. New model architectures vs. inference-engine capability is always a catch-up race — SGLang and vLLM spent half a year getting efficient DeepSeek-V3 inference working.
- The inference characteristics of SSM/Mamba: Mamba's "stateful inference" is completely different from attention's KV cache, and engines must be rewritten. The PagedAttention equivalent for SSM inference has not appeared yet — an open opportunity.
- The complexity of hybrid architectures: a Mamba+Transformer hybrid like Jamba forces an engine to support two inference paradigms at once, at high engineering cost. Short-term MoE remains mainstream; Mamba and Linear Attention are complements, not replacements. Architecture mechanics: GPU Architecture and Optimization and Distributed Inference (TP/PP).
Line 6: Hardware Generations — B200, MI300X, Ascend, Groq, Cerebras
Status. A hardware-generation shift means rewriting the algorithm. H100 introduced FP8, B200 introduces FP4 and NVLink 5.0; AMD's MI300X challenges NVIDIA with HBM3+ and large memory; Huawei's Ascend 910B carries a large share of domestic inference load; Groq's LPU pursues extreme throughput with a deterministic architecture; Cerebras' WSE computes on an entire wafer.
Representative work / hardware releases.
| Hardware | Key characteristics | Impact on inference |
|---|---|---|
| NVIDIA B200 / GB200 NVL72 (Blackwell, 2024) | FP4 Tensor Cores + NVFP4 format + 72-GPU NVLink cluster | FP4 quantization sweet spot; large models servable on a single cluster |
| NVIDIA H100 / H200 (Hopper, 2022-2024) | FP8 Tensor Cores + Transformer Engine + 4th-gen NVLink | FP8 inference sweet spot; the hardware basis of FA3 |
| AMD MI300X (2023) | 192 GB HBM3 + 5.3 TB/s bandwidth | Large memory + high bandwidth, challenging NVIDIA; supported by vLLM/SGLang |
| Huawei Ascend 910B (2023) | Domestic 7nm + Da Vinci architecture + 64 GB HBM | One of the main domestic hardware platforms for large-model inference |
| Groq LPU (commercial 2024) | Deterministic architecture + SRAM-dominant + extreme throughput | Very low per-token latency, but small memory, unsuited to large-model KV caches |
| Cerebras WSE-3 (2024) | Whole wafer + 4 trillion transistors | Extreme sparse compute, but weak ecosystem and complex deployment |
| Intel Gaudi 2/3 (2023-2024) | 96 GB HBM2 + 24 GB HBM3 | Cost-effective route, deployed by some enterprises |
Key questions.
- A hardware generation means rewriting the algorithm: FA1 is optimal on A100, FA3 on H100, and FP4 kernels are only sweet on B200. Change the hardware generation and the operator must be rewritten — the core pain point and the ongoing opportunity of this field.
- The ecosystem catch-up of domestic hardware: the Ascend 910B is close to the A100 in raw performance, but its software ecosystem (PyTorch integration, vLLM porting, operator libraries) lags 2-3 years. Domestic inference engines and operators are the key opportunity of the next 3-5 years.
- The fate of extreme architectures: Groq LPU's "deterministic + high throughput" is stunning in specific scenarios, but small memory limits general LLM inference; Cerebras WSE has astonishing compute but a high deployment bar. Extreme architectures either find a niche and survive, or get swallowed by the mainstream GPU ecosystem. Hardware background: Hardware Primer.
3. Which Problems Are Still Unsolved
Every line above has "key questions," but a few problems are cross-line, repeatedly mentioned, and far from solved. Facing them honestly matters more than chasing hot spots.
3.1 The Conflict Between Tail Latency and Fairness
LLM serving throughput and average latency are easy to measure, but P99/P999 tail latency is extremely hard to control — it depends on the load distribution, the scheduling policy, the KV cache hit rate, and hardware jitter. There is no "SOTA for tail latency" in the industry, only case studies of "under load X, with engine Y, achieving tail latency Z." DistServe/SARATHI all attack it, but none is a "general solution."
3.2 A "Multiplexing Theory" for Multi-Model Deployment
AlpaServe gave a statistical framework for "time-multiplexing across models," but best practices for multi-model deployment are still not settled — every team is groping on its own for "how many models on how many GPUs, with what batching policy." This is an open opportunity and a daily pain point in industry.
3.3 The Training-Time Internalization of Speculative Decoding
DeepSeek-V3 MTP opened the "learn to draft at training time" paradigm, but the training cost of MTP, its impact on the main model's accuracy, and the portability of drafting ability are all undecided. The key open question of 2025-2026.
3.4 Unifying Quantization and Sparsity
Quantization (make weights smaller) and sparsity (make compute fewer) are two independent optimization routes — can they cooperate? For example, can a sparse attention pattern cooperate with weight quantization? Almost no one studies this intersection; it is an open opportunity.
3.5 The Operator Ecosystem of Domestic Hardware
The Ascend 910B and other domestic hardware are not weak in compute, but the operator ecosystem is a "blank field." Domestic-hardware implementations of mainstream work like FA3, PagedAttention, and speculative decoding are a hard requirement and an opportunity for the next 3-5 years.
How to live with "unsolved"
Do not equate "unsolved" with "no opportunity." On the contrary: high-value output usually appears at the edge of a recognized hard problem. Tail latency spawned the prosperity of DistServe/SARATHI; multi-model multiplexing spawned AlpaServe; the training-time internalization of speculative decoding spawned DeepSeek-V3 MTP. Understand the problem list and you understand the opportunity list.
4. Advice for Readers: How to Choose a Direction to Go Deep
Facing six lines, the most common confusion is "which one to pick." Three principles:
Principle 1: Pick by "your irreplaceability," not by "heat." All six lines are hot, but your background decides which has the lowest entry cost and the deepest moat:
| Your background | Entry direction | Why |
|---|---|---|
| Strong algorithms/CUDA | Line 1 (operators) or Line 4 (quantization) | Operators and quantization are compute-dense areas; CUDA skill is an advantage |
| Strong systems/distributed | Line 2 (serving systems) or Line 5 (architecture innovation) | Serving systems and MoE inference are system-dense areas |
| Strong engineering/product | Line 3 (speculative decoding) or Line 6 (hardware) | Speculative decoding is the hottest to productionize; hardware adaptation has many domestic gaps |
| Strong algorithms/modeling | Line 5 (architecture innovation) | Mamba/MoE have large design space, won by insight rather than compute |
| Domestic-hardware background | The domestic part of Line 6 (hardware) | The domestic operator ecosystem is the key opportunity of the next 3-5 years |
| Not sure yet | First read the Paper Map and Classic Papers in Depth | Spend a month or two building the full picture before betting; do not rush |
Principle 2: Go deep in one direction "until you can independently ship a project," then expand horizontally. The biggest trap of the frontier is "read every paper and dabble in every direction." A more effective pattern:
- Pick one direction and closely read its classic papers (see Classic Papers in Depth) in chronological order — ten to twenty papers — and draw a technical-evolution diagram.
- Use open-source weights and engines to reproduce at least one representative work (even a scaled-down version), turning the paper's "how" into a capability in your hands.
- Participate in one active PR on vLLM, SGLang, or TensorRT-LLM (file an issue, fix a bug, reproduce an experiment), letting real feedback calibrate your understanding.
- Use the tools and leaderboards in Curated Resources to keep tracking this direction, and write periodic personal surveys.
Principle 3: Beware narratives; embrace numbers. Frontier writing is full of "3x speedup," "lossless," and "revolutionary" narratives, but they rarely survive next month's scrutiny. Your defense is evaluation and reproduction: for any "big breakthrough," first find the data on a benchmark site, then run a small experiment to verify. Judging "what will stay and what will ebb" is more reliable than predicting "the next hot spot."
A one-year timeline for going deep (reference)
If you decide to go deep in one direction, an executable one-year frame: months 1-3, closely read the direction's classic papers and reproduce at least one scaled-down experiment, building your own code base; months 4-6, independently finish a small project (reproduction + one improvement of your own) and write it up as a technical report; months 7-9, join an open-source engine or a public benchmark, letting community feedback calibrate your judgment; months 10-12, package the results into a presentable form (blog, report, competition results). After a year you will own not "many papers read" but a loop you can walk independently — "read paper -> run experiment -> form judgment" — which is worth more than any amount of knowledge.
A time budget for newcomers
If you can invest 10 hours a week: 5 hours of close reading and reproduction, 3 hours of tracking (arXiv + leaderboards + repo PRs), and 2 hours of notes and summaries. After three months you will noticeably feel that judgment matters far more than memory. The full learning-path plan is in Learning Paths: Three Routes and Start Here.
5. Further Reading
- Start Here — the entry to this site's papers section, with search and reading methods
- Paper Map — the positional relationships of the six lines in one map
- Classic Papers in Depth — the close-reading versions of the key papers cited here
- Reading Discipline & FAQ — high-frequency Q&A on paper reading and research directions
- What Is Inference Acceleration? — the conceptual foundation of inference deployment
- Case studies: vLLM and PagedAttention, TensorRT-LLM, Speculative Decoding and Medusa/EAGLE, Distributed Inference (TP/PP), Mobile Deployment
- Concept pages: Kernel Fusion and Custom Kernels, GPU Architecture and Optimization, Model Quantization Fundamentals, Model Serving and Orchestration
- Curated Resources — the full tool list needed to track the frontier
- Hardware Primer — background on GPUs, NPUs, and domestic hardware
References
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (arXiv:2407.08608)
- FlashInfer: An Efficient and Customizable Attention Engine for LLMs (arXiv:2401.12448)
- vAttention: Efficient Attention with Suffix Parallelism (arXiv:2405.04447)
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (arXiv:2401.09670)
- Splitwise: Efficient generative LLM inference using phase splitting (HPCA 2024)
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills (arXiv:2308.16369)
- SGLang: Efficient Execution of Structured Language Model Programs (arXiv:2312.07104)
- Mooncake: A KV-Centric Decoupled Architecture for LLM Serving (arXiv:2407.00079)
- AlpaServe: Scaling for Large-Scale Deep Learning Serving (OSDI 2023)
- EAGLE-3: Scaling up Inference Speed (arXiv:2503.01840)
- Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (arXiv:2401.10774)
- EAGLE-2: Faster Speculative Decoding with Dynamic Draft Trees (arXiv:2406.16858)
- DeepSeek-V3 Technical Report (arXiv:2412.19437)
- QuaRot: Outlier-free 4-bit Inference for Rotated LLMs (arXiv:2404.00456)
- SpinQuant: LLM Quantization with Learned Rotations (arXiv:2405.16406)
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)
- Mixtral of Experts (arXiv:2401.04088)
- Lightning Attention-2 (arXiv:2401.11461)
- NVIDIA H100 / H200 data-center GPUs
- NVIDIA Blackwell B200 / GB200 NVL72
- AMD Instinct MI300X
- Groq LPU
- Cerebras WSE-3
- arXiv (latest cs.LG / cs.DC submissions)
- Hugging Face Papers
- Papers with Code
- Artificial Analysis
- vLLM benchmarks
- SGLang project
- FlashAttention official repo
- FlashInfer project
- TensorRT-LLM project