Appearance
Curated Resources
Inference acceleration is a hard-nosed discipline — one foot in PyTorch/CUDA engineering practice, one foot in systems papers like vLLM/PagedAttention, and a hand in hardware architecture. Over the past three years (2022–2026), the field has accumulated a set of extremely high-quality, mostly free official resources: the vLLM source and blog, the FlashAttention paper series, the CMU 15-819 ASPIRE course, the CUDA Mode podcast. They form a clear progression of skills and a division of styles. This page organizes them into seven topic groups, each entry with a one-line positioning, who it's for, and the official entry point — every link individually verified.
How to use
This list is not a "favorites folder" — it's an external resource layer paired with this site's Learning Paths: Three Routes. The recommended pattern: build a conceptual skeleton here first (starting from What Is Inference Acceleration? and the Glossary), then pick one or two main lines from this page and follow them to the end — greed spoils the meal: one line walked to the end beats ten that each stop at chapter three.
1. Official Documentation
Official docs for inference engines and low-level libraries are the authoritative "look it up as you go" source — more accurate than any secondhand tutorial. This section keeps only the core entry points you'll inevitably touch while learning inference.
vLLM Documentation
One-line positioning: the de facto standard of open-source LLM inference engines; the docs cover the full principles and APIs of PagedAttention, continuous batching, and prefix caching.
The vLLM docs come in three blocks: Getting Started (up and running with one line — pip install vllm), Design (the PagedAttention paper + architecture diagrams — a must-read for newcomers), and API Reference (all parameters of the offline LLM class and the online AsyncLLMEngine). The source itself is also a textbook: vllm/core/scheduler.py and vllm/core/block_manager.py are the best live sites for understanding continuous batching. See vLLM and PagedAttention.
- Who it's for: engineers deploying/tuning LLM inference services; researchers who want to understand the PagedAttention implementation.
- Official entry: docs.vllm.ai
TensorRT-LLM Documentation
One-line positioning: NVIDIA's official LLM inference engine, fusing FP8, in-flight batching, and speculative decoding deeply with Hopper/Blackwell hardware.
The TensorRT-LLM docs cover the three-stage Builder/Engine/Runtime pipeline, precision and quantization (FP8, INT4 AWQ, SmoothQuant), in-flight batching, speculative decoding (Medusa/EAGLE), and Multi-Device (TP/PP) configuration. The companion TensorRT-LLM benchmark repo is a first-hand reference for performance tuning. See TensorRT-LLM and TensorRT and GPU Inference.
- Who it's for: engineers squeezing the last drop of performance on H100/H200/B200; anyone deploying FP8/FP4.
- Official entry: nvidia.github.io/TensorRT-LLM
NVIDIA Triton Inference Server Documentation
One-line positioning: NVIDIA's multi-backend inference serving framework — unified management of TensorRT/PyTorch/ONNX/TF models with HTTP/gRPC and dynamic batching.
The Triton docs cover model repository management (repository layout), backends (TensorRT, PyTorch, ONNX Runtime, Python backend), dynamic batching policies, and the Model Analyzer. It is the de facto way to wrap multiple engines into one unified service. See Triton Inference Server.
- Who it's for: anyone unifying multiple models/engines into one online service; anyone doing model version management and canary rollout.
- Official entry: docs.nvidia.com/deeplearning/triton-inference-server
The two Tritons
The official docs at triton-inference-server are the serving framework; triton-lang is OpenAI's GPU kernel DSL. Same name, unrelated things — this section means the former; for the latter see the chapters around Kernel Fusion and Custom Kernels.
ONNX Runtime Documentation
One-line positioning: Microsoft's cross-platform inference engine covering CPU/GPU/NPU/mobile — the engineering base for "one model, run it anywhere."
The ONNX Runtime docs are organized into Execution Providers (CUDA, TensorRT, OpenVINO, CoreML, DirectML and other backends), Graph Optimizer (graph-optimization rules), and Performance Tuning (thread affinity, memory patterns). Its cross-platform nature is key for edge and hybrid deployment. See ONNX Runtime: Cross-Platform.
- Who it's for: anyone deploying the same model across hardware; CPU/mobile inference engineers.
- Official entry: onnxruntime.ai
OpenVINO Documentation
One-line positioning: Intel's inference optimization toolchain for CPU/iGPU/VPU — squeezes matmuls to the limit on Intel hardware (including AMX).
The OpenVINO docs cover the Model Optimizer (converting PyTorch/TF models to IR), NNCF (the Neural Network Compression Framework — quantization/pruning/distillation), and the Inference Engine runtime. It is the first choice for Intel CPU inference. See OpenVINO and CPU Inference.
- Who it's for: anyone deploying inference on Intel CPU/iGPU; anyone wringing matmul performance out of CPUs with AMX.
- Official entry: docs.openvino.ai
CUDA Toolkit / cuDNN / cuBLAS Documentation
One-line positioning: the official docs for NVIDIA GPU programming and operator libraries — the foundation for understanding GPU microarchitecture and kernel optimization.
The CUDA Toolkit includes the CUDA C++ Programming Guide (SIMT, warps, shared memory, streams), the CUDA C++ Best Practices Guide (optimization techniques), and the Nsight docs. cuDNN/cuBLAS document the underlying libraries for GEMM/convolution/normalization. See GPU Architecture and Optimization and Hardware Primer.
- Who it's for: anyone writing custom CUDA kernels; anyone doing kernel-level performance tuning.
- Official entry: docs.nvidia.com/cuda
PyTorch Documentation
One-line positioning: the mainstream framework for training and inference — and since 2.x, torch.compile is the entry point into graph compilation.
The PyTorch docs' Tutorials include must-reads like "from eager to torch.compile" and "Transformer from scratch"; the torch.compile docs are the hands-on entry to JIT compilation and kernel fusion; the Distributed docs (DDP, FSDP, TP) are the foundation of large-model inference. See Inference vs. Training vs. Fine-Tuning and Distributed Inference (TP/PP).
- Who it's for: anyone moving PyTorch models from training to inference; anyone who wants to understand graph compilation behind
torch.compile. - Official entry: pytorch.org/docs/stable
How to read official docs
Framework docs are not textbooks — don't read them cover to cover. The right pattern: spend 30 minutes on Getting Started to build a mental model, enter the API pages with a specific question ("what does this parameter mean"), and finally use the example pages as templates for your own code. For trade-offs between engines, see Inference Engine Comparison.
2. Papers (arXiv)
Nearly every key technique in inference acceleration has a milestone paper. Below are must-reads by topic, each with a one-line positioning; for full close reading see Classic Papers in Depth and Paper Map.
Attention acceleration
- FlashAttention v1 (arxiv 2205.14135): tiling + online softmax avoids materializing the N×N matrix — 2–4× attention speedup. See Kernel Fusion and Custom Kernels.
- FlashAttention v2 (arxiv 2307.08691): fewer non-MMA instructions and better parallelism partitioning — 2× over v1 on A100.
- FlashAttention v3 (arxiv 2407.08608): Hopper-specialized, FP8 + asynchronous TMA — near-peak bandwidth on H100.
- FlashInfer (arxiv 2401.02081): an attention kernel library purpose-built for LLM inference, covering prefill/decode/append.
- PagedAttention (arxiv 2309.06180): vLLM's paged KV cache management — OS virtual memory ideas moved into LLMs. See vLLM and PagedAttention.
Quantization
- GPTQ (arxiv 2210.17323): second-order-Hessian PTQ with per-column error compensation — one of the de facto INT4 LLM standards. See Weight-Only Quantization and Mixed Precision.
- AWQ (arxiv 2306.00978): weight-only quantization that protects "important channels" — the other INT4 mainstream.
- SmoothQuant (arxiv 2211.10438): migrates hard-to-quantize activation outliers to the weight side, achieving W8A8 INT8.
- ZeroQuant (arxiv 2206.01861): ByteDance's per-token + per-group quantization scheme.
- FP8 paper and spec: FP8 Formats for Deep Learning, jointly published by NVIDIA, Intel, and AMD — the foundation of Hopper's native precision.
Speculative decoding and parallelism
- Medusa (arxiv 2401.10774): multiple heads guess tokens in parallel and the original model verifies in one pass — no draft model needed. See Speculative Decoding and Medusa/EAGLE.
- EAGLE-2/3 (arxiv 2401.15077, arxiv 2503.01840): autoregressive draft + dynamic tree — 2–3× throughput.
- DistServe (arxiv 2401.09670): disaggregates prefill and decode onto different GPUs, each optimized separately. See Batching and Request Scheduling.
- Splitwise (arxiv 2311.18677): a similar prefill/decode disaggregation architecture.
- SARATHI (arxiv 2308.16369): chunked prefill — slicing long prompts to be co-scheduled with decode.
- DeepSeek-V3 technical report (arxiv 2412.19437): MoE training and inference engineering practice — the real-world case of EP parallelism.
Engines and systems
- vLLM (arxiv 2309.06180): the original paper of PagedAttention + continuous batching.
- SGLang (arxiv 2312.07104): RadixAttention prefix caching and the structured-generation programming model.
- DeepSpeed-FastGen (arxiv 2401.08671): Microsoft's LLM inference system — the Dynamic SplitFuse idea.
More papers organized by reading path are in Reading Paths and Paper Map; frontier progress in Frontier Advances.
3. Open-Source Repositories
Reading source code is the fastest path into an inference engine's internals. This section is organized "from mainstream to supporting."
Mainstream LLM inference engines
- vllm-project/vllm: the de facto open-source LLM inference engine, started from PagedAttention + continuous batching. See vLLM and PagedAttention.
- NVIDIA/TensorRT-LLM: NVIDIA's official LLM engine, deeply fused with FP8 hardware. See TensorRT-LLM.
- sglang-project/SGLang: structured generation and RadixAttention prefix caching.
- huggingface/text-generation-inference (TGI): HF's official LLM serving engine.
- ggerganov/llama.cpp: a C++ all-platform LLM inference library — GGUF is the de facto edge standard. See llama.cpp and GGUF.
Classic inference engines and compilers
- NVIDIA/TensorRT: NVIDIA's GPU inference optimization engine — the foundation under TensorRT-LLM. See TensorRT and GPU Inference.
- microsoft/onnxruntime: the cross-platform inference engine. See ONNX Runtime: Cross-Platform.
- NVIDIA/triton-inference-server: the multi-backend inference serving framework. See Triton Inference Server.
- apache/tvm: the end-to-end deep learning compiler — graph optimization + kernel auto-tuning. See Computation Graph Optimization.
Kernels and low-level
- openai/triton: the GPU kernel DSL — FlashAttention v2 onward is built on it. See Kernel Fusion and Custom Kernels.
- flashinfer-ai/flashinfer: the LLM attention kernel library adopted by vLLM/SGLang.
- Dao-AILab/flash-attention: the official FlashAttention implementation, covering v1/v2/v3.
- comefromplutozh/w8a8 (community implementation): a teaching reference for SmoothQuant and W8A8 quantization.
Quantization and compression
- AUTOMATIC1111/stable-diffusion-webui: community quantization deployment recipes for Stable Diffusion — INT4/INT8 practice reference.
- casper-hansen/AutoAWQ: the official AWQ implementation.
- IST-DASLab/gptq: the official GPTQ implementation.
4. Blogs and Tutorials
To keep pace with inference acceleration, blogs are more timely than papers and easier to understand than official docs. This section keeps the highest-quality ones.
Engine official blogs
- vLLM blog: the official publication venue for PagedAttention, continuous batching, prefix caching, and SGLang comparisons — every post ships measured data.
- SGLang blog: official write-ups on RadixAttention, structured generation, and multi-turn dialogue optimization.
- NVIDIA Developer Blog: deep technical dives on TensorRT-LLM, FP8, and the Hopper Transformer Engine.
Personal deep-dive blogs
- Jay Shah's blog: line-by-line source-level walkthroughs of PagedAttention, continuous batching, and speculative decoding — recommended by the vLLM team.
- Aleksa's "attention" series: the evolution from the plainest attention to FlashAttention v3, with visualizations.
- Lilian Weng — Lil'Log: long-form posts by a former OpenAI research director — the quantization and LLM inference systems series is extremely dense.
- Hugging LLM Course: HF's official LLM course, with inference and deployment chapters.
- Soumith Chintala's blog and talks: one of PyTorch's creators — insights on GPUs and deep learning systems.
Podcasts and video
- GPU Mode Discord: the community for GPU programming and inference optimization — talks every week, extremely high signal.
- CUDA Mode YouTube: GPU Mode's video channel — hands-on FlashAttention, Triton, and Nsight walkthroughs.
- Karpathy YouTube: mostly about training, but "Let's build GPT" and "intro to LLMs" are hugely helpful for understanding inference mechanics.
5. Courses
Academic courses are the best path to a systematic "inference optimization mental model." Below are the recognized top-tier ones — all free.
CMU 15-819 ASPIRE (Hazy Research)
One-line positioning: Tri Dao of Stanford Hazy Research teaches "Efficient Deep Learning" at CMU — first-hand lectures on FlashAttention and systems optimization.
Covers attention optimization, kernel fusion, long context, sparse attention, and sampling acceleration. Tri Dao is FlashAttention's first author, making this a rare case of "the author teaching his own work." The companion notes and the Hazy Research blog complement each other.
- Who it's for: anyone going deep on FlashAttention and long-context optimization; inference engine developers.
- Official entry: cs.cmu.edu/~15319
MIT 6.5940 (Song Han's team)
One-line positioning: Han Lab's "TinyML and Efficient Deep Learning" — a systematic treatment of quantization, pruning, and distillation.
Song Han is one of the founders of deep learning compression (Deep Compression), and this course covers the full chain from algorithms to hardware — INT4/INT8 quantization, structured pruning, knowledge distillation — with lab assignments. See Model Quantization Fundamentals, Pruning and Sparsification, and Knowledge Distillation.
- Who it's for: anyone doing model compression and edge deployment; anyone who wants the principles of quantization/pruning.
- Official entry: hanlab.mit.edu/courses/2024-fall-65940
Stanford CS336
One-line positioning: Stanford's "build a large language model from scratch" course — covering training + inference + systems.
CS336 doesn't stop at training LLMs; it teaches efficient inference — hands-on KV cache, attention optimization, and quantization. Assignments require writing an attention kernel and implementing KV cache from zero. It is the best single course for the "training-to-inference full stack."
- Who it's for: anyone building the full LLM systems stack; anyone who wants to implement KV cache and attention from scratch.
- Official entry: stanford-cs336.github.io/s2025
Sasha Rush's annotated tensor series
One-line positioning: Cornell professor Sasha Rush's "annotated tensor" series — walking through attention, RoPE, and quantization line by line in Jupyter.
Sasha Rush (one of the authors of PyTorch TorchDynamo) turns abstract concepts into runnable code with annotated notebooks — from attention's matrix multiplications to FlashAttention's tiling to the dequantize GEMM of INT4 quantization, every step paired with code and diagrams. The best "learn by doing" resource.
- Who it's for: people who learn from code rather than formulas; anyone who wants to implement attention/quantization from zero.
- Official entry: srush.github.io/tensor-loaded
Course pairing suggestions
Start with CS336 to build the whole-system view of LLM systems (including inference); go deep on quantization/pruning with MIT 6.5940; go deep on attention optimization and long context with CMU 15-819; pair everything with Sasha Rush's notebooks, reading and coding along. Interleave with this site's Learning Paths: Three Routes as needed.
6. Benchmarks and Tools
Benchmark leaderboards
- MLPerf Inference: the industry's most authoritative inference benchmark, including Llama-2-70B and Llama-3.1-405B. See Benchmark Data & Tool Profiles.
- Hugging Face Open LLM Leaderboard: the main model-quality leaderboard — the first reference for model selection.
- Papers with Code: papers + code + benchmarks in one place — find SOTA and its implementation.
Evaluation tools
- lm-evaluation-harness: EleutherAI's LLM quality evaluation framework — the engine under the Open LLM Leaderboard.
- vLLM benchmark suite: throughput and latency stress-testing for LLM serving.
- sglang bench: stress-testing for structured generation and multi-turn dialogue.
Profiling tools
- Nsight Systems: GPU timelines and kernel time shares — the first tool for diagnosing inference bottlenecks. See Tuning and Performance Optimization.
- Nsight Compute: per-kernel compute/bandwidth/occupancy analysis.
- PyTorch Profiler: operator-level time and memory in PyTorch.
- Triton tutorial: the entry course for writing custom kernels, from vector add to fused attention.
7. Communities
Keeping pace with the field relies on fixed information inlets. The communities below cover discussion, Q&A, and news.
Discord / Slack
- GPU Mode Discord: the most active home for GPU programming and inference optimization — authors answer questions in person. Strongly recommended.
- vLLM Slack: the official vLLM community — first stop for bug reports and feature requests.
- NVIDIA Developer Forums: official Q&A for TensorRT, CUDA, and cuDNN — questions and answers are searchable and stay on record.
- Hugging Face Discord: discussion across the HF ecosystem — Transformers/TGI questions welcome.
Chinese-language communities
- Synced (Jiqizhixin): academically and industrially deep — high-quality paper interpretations.
- QbitAI: industry dynamics and productization news — fast.
- Zhihu's "LLM inference" topic: the highest-quality Chinese-language discussion board for LLM inference, with lots of long-form practitioner writing. Verify timeliness before trusting specific claims.
The golden ratio of information intake
Daily: scan arXiv cs.LG/cs.DC titles (15 minutes); weekly: watch GPU Mode Discord's featured talks + read one newsletter; monthly: closely read one key paper from this page. Treat news accounts as supplements, never primary sources — capability comes from papers and code, not from news.
How to Chain Them Together
Finally, a "stage → main-line resources → site companions" table — take what you need:
| Your stage | Main-line resources (this page) | Site companions |
|---|---|---|
| New to inference optimization, not sure what it is | vLLM docs Getting Started + What Is Inference Acceleration? | Glossary |
| Can run vLLM, want the principles | vLLM blog + PagedAttention paper + Jay Shah's blog | vLLM and PagedAttention |
| Need to squeeze H100 dry | TensorRT-LLM docs + NVIDIA blog + Nsight tools | TensorRT-LLM + Tuning and Performance Optimization |
| Want to understand quantization | MIT 6.5940 + AWQ/GPTQ papers + Sasha Rush's notebooks | Model Quantization Fundamentals + Weight-Only Quantization and Mixed Precision |
| Want to understand attention optimization | CMU 15-819 + FlashAttention v1/v2/v3 + FlashInfer | Kernel Fusion and Custom Kernels |
| Doing edge deployment | llama.cpp + ONNX Runtime Mobile + MIT 6.5940 | llama.cpp and GGUF + Mobile Deployment |
| Doing distributed inference | NCCL docs + DeepSpeed-FastGen + DistServe paper | Distributed Inference (TP/PP) |
| Want to benchmark yourself | MLPerf + vLLM benchmark suite + Nsight Systems | Benchmark Data & Tool Profiles + Inference Benchmarking in Practice |
Remember three iron rules: pick one main line; doing beats collecting; official docs and papers beat secondhand tutorials. A hundred saved links are worth less than one project that runs — now close this page and pick a line from Learning Paths: Three Routes to start.