Skip to content

Curated Resources

At a glance From vLLM to TensorRT-LLM, from FlashAttention to PagedAttention, from CUDA Mode to CMU 15-819 — a topic-organized map of learning resources for LLM inference acceleration and deployment, with every official link verified.

Curated Resources ​

Inference acceleration is a hard-nosed discipline — one foot in PyTorch/CUDA engineering practice, one foot in systems papers like vLLM/PagedAttention, and a hand in hardware architecture. Over the past three years (2022–2026), the field has accumulated a set of extremely high-quality, mostly free official resources: the vLLM source and blog, the FlashAttention paper series, the CMU 15-819 ASPIRE course, the CUDA Mode podcast. They form a clear progression of skills and a division of styles. This page organizes them into seven topic groups, each entry with a one-line positioning, who it's for, and the official entry point — every link individually verified.

How to use

This list is not a "favorites folder" — it's an external resource layer paired with this site's Learning Paths: Three Routes. The recommended pattern: build a conceptual skeleton here first (starting from What Is Inference Acceleration? and the Glossary), then pick one or two main lines from this page and follow them to the end — greed spoils the meal: one line walked to the end beats ten that each stop at chapter three.

1. Official Documentation ​

Official docs for inference engines and low-level libraries are the authoritative "look it up as you go" source — more accurate than any secondhand tutorial. This section keeps only the core entry points you'll inevitably touch while learning inference.

vLLM Documentation ​

One-line positioning: the de facto standard of open-source LLM inference engines; the docs cover the full principles and APIs of PagedAttention, continuous batching, and prefix caching.

The vLLM docs come in three blocks: Getting Started (up and running with one line — pip install vllm), Design (the PagedAttention paper + architecture diagrams — a must-read for newcomers), and API Reference (all parameters of the offline LLM class and the online AsyncLLMEngine). The source itself is also a textbook: vllm/core/scheduler.py and vllm/core/block_manager.py are the best live sites for understanding continuous batching. See vLLM and PagedAttention.

  • Who it's for: engineers deploying/tuning LLM inference services; researchers who want to understand the PagedAttention implementation.
  • Official entry: docs.vllm.ai

TensorRT-LLM Documentation ​

One-line positioning: NVIDIA's official LLM inference engine, fusing FP8, in-flight batching, and speculative decoding deeply with Hopper/Blackwell hardware.

The TensorRT-LLM docs cover the three-stage Builder/Engine/Runtime pipeline, precision and quantization (FP8, INT4 AWQ, SmoothQuant), in-flight batching, speculative decoding (Medusa/EAGLE), and Multi-Device (TP/PP) configuration. The companion TensorRT-LLM benchmark repo is a first-hand reference for performance tuning. See TensorRT-LLM and TensorRT and GPU Inference.

  • Who it's for: engineers squeezing the last drop of performance on H100/H200/B200; anyone deploying FP8/FP4.
  • Official entry: nvidia.github.io/TensorRT-LLM

NVIDIA Triton Inference Server Documentation ​

One-line positioning: NVIDIA's multi-backend inference serving framework — unified management of TensorRT/PyTorch/ONNX/TF models with HTTP/gRPC and dynamic batching.

The Triton docs cover model repository management (repository layout), backends (TensorRT, PyTorch, ONNX Runtime, Python backend), dynamic batching policies, and the Model Analyzer. It is the de facto way to wrap multiple engines into one unified service. See Triton Inference Server.

The two Tritons

The official docs at triton-inference-server are the serving framework; triton-lang is OpenAI's GPU kernel DSL. Same name, unrelated things — this section means the former; for the latter see the chapters around Kernel Fusion and Custom Kernels.

ONNX Runtime Documentation ​

One-line positioning: Microsoft's cross-platform inference engine covering CPU/GPU/NPU/mobile — the engineering base for "one model, run it anywhere."

The ONNX Runtime docs are organized into Execution Providers (CUDA, TensorRT, OpenVINO, CoreML, DirectML and other backends), Graph Optimizer (graph-optimization rules), and Performance Tuning (thread affinity, memory patterns). Its cross-platform nature is key for edge and hybrid deployment. See ONNX Runtime: Cross-Platform.

  • Who it's for: anyone deploying the same model across hardware; CPU/mobile inference engineers.
  • Official entry: onnxruntime.ai

OpenVINO Documentation ​

One-line positioning: Intel's inference optimization toolchain for CPU/iGPU/VPU — squeezes matmuls to the limit on Intel hardware (including AMX).

The OpenVINO docs cover the Model Optimizer (converting PyTorch/TF models to IR), NNCF (the Neural Network Compression Framework — quantization/pruning/distillation), and the Inference Engine runtime. It is the first choice for Intel CPU inference. See OpenVINO and CPU Inference.

  • Who it's for: anyone deploying inference on Intel CPU/iGPU; anyone wringing matmul performance out of CPUs with AMX.
  • Official entry: docs.openvino.ai

CUDA Toolkit / cuDNN / cuBLAS Documentation ​

One-line positioning: the official docs for NVIDIA GPU programming and operator libraries — the foundation for understanding GPU microarchitecture and kernel optimization.

The CUDA Toolkit includes the CUDA C++ Programming Guide (SIMT, warps, shared memory, streams), the CUDA C++ Best Practices Guide (optimization techniques), and the Nsight docs. cuDNN/cuBLAS document the underlying libraries for GEMM/convolution/normalization. See GPU Architecture and Optimization and Hardware Primer.

  • Who it's for: anyone writing custom CUDA kernels; anyone doing kernel-level performance tuning.
  • Official entry: docs.nvidia.com/cuda

PyTorch Documentation ​

One-line positioning: the mainstream framework for training and inference — and since 2.x, torch.compile is the entry point into graph compilation.

The PyTorch docs' Tutorials include must-reads like "from eager to torch.compile" and "Transformer from scratch"; the torch.compile docs are the hands-on entry to JIT compilation and kernel fusion; the Distributed docs (DDP, FSDP, TP) are the foundation of large-model inference. See Inference vs. Training vs. Fine-Tuning and Distributed Inference (TP/PP).

  • Who it's for: anyone moving PyTorch models from training to inference; anyone who wants to understand graph compilation behind torch.compile.
  • Official entry: pytorch.org/docs/stable

How to read official docs

Framework docs are not textbooks — don't read them cover to cover. The right pattern: spend 30 minutes on Getting Started to build a mental model, enter the API pages with a specific question ("what does this parameter mean"), and finally use the example pages as templates for your own code. For trade-offs between engines, see Inference Engine Comparison.

2. Papers (arXiv) ​

Nearly every key technique in inference acceleration has a milestone paper. Below are must-reads by topic, each with a one-line positioning; for full close reading see Classic Papers in Depth and Paper Map.

Attention acceleration ​

Quantization ​

Speculative decoding and parallelism ​

Engines and systems ​

  • vLLM (arxiv 2309.06180): the original paper of PagedAttention + continuous batching.
  • SGLang (arxiv 2312.07104): RadixAttention prefix caching and the structured-generation programming model.
  • DeepSpeed-FastGen (arxiv 2401.08671): Microsoft's LLM inference system — the Dynamic SplitFuse idea.

More papers organized by reading path are in Reading Paths and Paper Map; frontier progress in Frontier Advances.

3. Open-Source Repositories ​

Reading source code is the fastest path into an inference engine's internals. This section is organized "from mainstream to supporting."

Mainstream LLM inference engines ​

Classic inference engines and compilers ​

Kernels and low-level ​

Quantization and compression ​

4. Blogs and Tutorials ​

To keep pace with inference acceleration, blogs are more timely than papers and easier to understand than official docs. This section keeps the highest-quality ones.

Engine official blogs ​

  • vLLM blog: the official publication venue for PagedAttention, continuous batching, prefix caching, and SGLang comparisons — every post ships measured data.
  • SGLang blog: official write-ups on RadixAttention, structured generation, and multi-turn dialogue optimization.
  • NVIDIA Developer Blog: deep technical dives on TensorRT-LLM, FP8, and the Hopper Transformer Engine.

Personal deep-dive blogs ​

  • Jay Shah's blog: line-by-line source-level walkthroughs of PagedAttention, continuous batching, and speculative decoding — recommended by the vLLM team.
  • Aleksa's "attention" series: the evolution from the plainest attention to FlashAttention v3, with visualizations.
  • Lilian Weng — Lil'Log: long-form posts by a former OpenAI research director — the quantization and LLM inference systems series is extremely dense.
  • Hugging LLM Course: HF's official LLM course, with inference and deployment chapters.
  • Soumith Chintala's blog and talks: one of PyTorch's creators — insights on GPUs and deep learning systems.

Podcasts and video ​

  • GPU Mode Discord: the community for GPU programming and inference optimization — talks every week, extremely high signal.
  • CUDA Mode YouTube: GPU Mode's video channel — hands-on FlashAttention, Triton, and Nsight walkthroughs.
  • Karpathy YouTube: mostly about training, but "Let's build GPT" and "intro to LLMs" are hugely helpful for understanding inference mechanics.

5. Courses ​

Academic courses are the best path to a systematic "inference optimization mental model." Below are the recognized top-tier ones — all free.

CMU 15-819 ASPIRE (Hazy Research) ​

One-line positioning: Tri Dao of Stanford Hazy Research teaches "Efficient Deep Learning" at CMU — first-hand lectures on FlashAttention and systems optimization.

Covers attention optimization, kernel fusion, long context, sparse attention, and sampling acceleration. Tri Dao is FlashAttention's first author, making this a rare case of "the author teaching his own work." The companion notes and the Hazy Research blog complement each other.

  • Who it's for: anyone going deep on FlashAttention and long-context optimization; inference engine developers.
  • Official entry: cs.cmu.edu/~15319

MIT 6.5940 (Song Han's team) ​

One-line positioning: Han Lab's "TinyML and Efficient Deep Learning" — a systematic treatment of quantization, pruning, and distillation.

Song Han is one of the founders of deep learning compression (Deep Compression), and this course covers the full chain from algorithms to hardware — INT4/INT8 quantization, structured pruning, knowledge distillation — with lab assignments. See Model Quantization Fundamentals, Pruning and Sparsification, and Knowledge Distillation.

Stanford CS336 ​

One-line positioning: Stanford's "build a large language model from scratch" course — covering training + inference + systems.

CS336 doesn't stop at training LLMs; it teaches efficient inference — hands-on KV cache, attention optimization, and quantization. Assignments require writing an attention kernel and implementing KV cache from zero. It is the best single course for the "training-to-inference full stack."

  • Who it's for: anyone building the full LLM systems stack; anyone who wants to implement KV cache and attention from scratch.
  • Official entry: stanford-cs336.github.io/s2025

Sasha Rush's annotated tensor series ​

One-line positioning: Cornell professor Sasha Rush's "annotated tensor" series — walking through attention, RoPE, and quantization line by line in Jupyter.

Sasha Rush (one of the authors of PyTorch TorchDynamo) turns abstract concepts into runnable code with annotated notebooks — from attention's matrix multiplications to FlashAttention's tiling to the dequantize GEMM of INT4 quantization, every step paired with code and diagrams. The best "learn by doing" resource.

  • Who it's for: people who learn from code rather than formulas; anyone who wants to implement attention/quantization from zero.
  • Official entry: srush.github.io/tensor-loaded

Course pairing suggestions

Start with CS336 to build the whole-system view of LLM systems (including inference); go deep on quantization/pruning with MIT 6.5940; go deep on attention optimization and long context with CMU 15-819; pair everything with Sasha Rush's notebooks, reading and coding along. Interleave with this site's Learning Paths: Three Routes as needed.

6. Benchmarks and Tools ​

Benchmark leaderboards ​

Evaluation tools ​

  • lm-evaluation-harness: EleutherAI's LLM quality evaluation framework — the engine under the Open LLM Leaderboard.
  • vLLM benchmark suite: throughput and latency stress-testing for LLM serving.
  • sglang bench: stress-testing for structured generation and multi-turn dialogue.

Profiling tools ​

7. Communities ​

Keeping pace with the field relies on fixed information inlets. The communities below cover discussion, Q&A, and news.

Discord / Slack ​

  • GPU Mode Discord: the most active home for GPU programming and inference optimization — authors answer questions in person. Strongly recommended.
  • vLLM Slack: the official vLLM community — first stop for bug reports and feature requests.
  • NVIDIA Developer Forums: official Q&A for TensorRT, CUDA, and cuDNN — questions and answers are searchable and stay on record.
  • Hugging Face Discord: discussion across the HF ecosystem — Transformers/TGI questions welcome.

Chinese-language communities ​

  • Synced (Jiqizhixin): academically and industrially deep — high-quality paper interpretations.
  • QbitAI: industry dynamics and productization news — fast.
  • Zhihu's "LLM inference" topic: the highest-quality Chinese-language discussion board for LLM inference, with lots of long-form practitioner writing. Verify timeliness before trusting specific claims.

The golden ratio of information intake

Daily: scan arXiv cs.LG/cs.DC titles (15 minutes); weekly: watch GPU Mode Discord's featured talks + read one newsletter; monthly: closely read one key paper from this page. Treat news accounts as supplements, never primary sources — capability comes from papers and code, not from news.

How to Chain Them Together ​

Finally, a "stage → main-line resources → site companions" table — take what you need:

Your stageMain-line resources (this page)Site companions
New to inference optimization, not sure what it isvLLM docs Getting Started + What Is Inference Acceleration?Glossary
Can run vLLM, want the principlesvLLM blog + PagedAttention paper + Jay Shah's blogvLLM and PagedAttention
Need to squeeze H100 dryTensorRT-LLM docs + NVIDIA blog + Nsight toolsTensorRT-LLM + Tuning and Performance Optimization
Want to understand quantizationMIT 6.5940 + AWQ/GPTQ papers + Sasha Rush's notebooksModel Quantization Fundamentals + Weight-Only Quantization and Mixed Precision
Want to understand attention optimizationCMU 15-819 + FlashAttention v1/v2/v3 + FlashInferKernel Fusion and Custom Kernels
Doing edge deploymentllama.cpp + ONNX Runtime Mobile + MIT 6.5940llama.cpp and GGUF + Mobile Deployment
Doing distributed inferenceNCCL docs + DeepSpeed-FastGen + DistServe paperDistributed Inference (TP/PP)
Want to benchmark yourselfMLPerf + vLLM benchmark suite + Nsight SystemsBenchmark Data & Tool Profiles + Inference Benchmarking in Practice

Remember three iron rules: pick one main line; doing beats collecting; official docs and papers beat secondhand tutorials. A hundred saved links are worth less than one project that runs — now close this page and pick a line from Learning Paths: Three Routes to start.