A Brief History
From 2010s CPU inference and OpenVINO, to the rise of GPU/cuDNN, to the Transformer and LLM inference revolution, and on to vLLM/PagedAttention, FP8, and speculative decoding. This page traces fifteen years of inference acceleration on a single timeline and the technical reasons behind each leap.