Appearance
Glossary
This glossary collects the core terminology of inference acceleration, organized into seven groups: fundamentals, performance & metrics, memory & compute, quantization & compression, operators & graphs, systems & serving, and hardware. Each entry gives the canonical English term and a precise definition; easily confused term pairs get their own side-by-side treatment. It is the dictionary you check while reading the rest of the site, cross-referenced with the Hardware Primer and Benchmark Data & Tool Profiles.
How to use
Don't read the terms in order. Skim "Fundamentals" first to build a vocabulary skeleton, then come back whenever you hit an unfamiliar word in other chapters. Bolded in-site links point to the deep-dive page for that term — for example, if you want to go further on quantization, jump to Model Quantization Fundamentals.
Fundamentals
inference The process of feeding input into a trained model and getting output — the core topic of this site. See What Is Inference Acceleration? and Inference vs. Training vs. Fine-Tuning.
training Iteratively updating model parameters with data so the model learns patterns from it. Training and inference share the same math (matrix multiplication, backpropagation) but pursue opposite goals — training changes parameters; inference uses them. See Inference vs. Training vs. Fine-Tuning.
fine-tuning Continuing to train a pretrained model on task data with a small learning rate. The relationship with inference: fine-tuning produces model weights; inference consumes them.
deployment The process of wrapping a trained model into a stable, business-callable service — model format conversion, engine selection, serving-ization, monitoring, and more. See Model Serving and Orchestration and Deployment Design Principles.
inference engine A runtime specialized in executing model inference — faster and more memory-efficient than general-purpose training frameworks. Representatives: TensorRT, ONNX Runtime, vLLM, Triton Inference Server. See Inference Engine Comparison.
serving The capability of wrapping a model behind HTTP/gRPC interfaces, with batching, load balancing, and version management. See Model Serving and Orchestration.
AOT and JIT (ahead-of-time / just-in-time) Two compilation timings. AOT compiles the whole graph into an immutable execution plan before deployment (TensorRT's Plan files); JIT compiles at runtime (PyTorch eager + torch.compile). AOT starts slowly but runs fast; JIT is flexible but carries overhead.
optimization The umbrella term for everything that makes inference better under given constraints (accuracy, latency, memory, cost): quantization, fusion, scheduling, and parallelism are all its subsets. See Tuning and Performance Optimization.
quantization Compressing model weights/activations from FP32 down to FP16/INT8/INT4/FP8/FP4 — halving memory, accelerating inference, with controllable accuracy loss. The first choice for deployment optimization. See Model Quantization Fundamentals.
pruning Zeroing out unimportant weights (structured pruning deletes channels/layers outright), shrinking the model; sparsity can then accelerate inference. See Pruning and Sparsification.
distillation Teaching a small model (the student) with the outputs of a large model (the teacher), so the student approximates the large model's quality at lower cost. See Knowledge Distillation.
graph optimization Rewrites at the computation-graph level: constant folding, operator fusion, layout conversion, and friends. See Computation Graph Optimization.
Performance & Metrics
latency The time from submitting one request to its completion. LLM inference breaks it into three points: TTFT, TPOT, and E2E. See Latency, Throughput, and Concurrency.
throughput Requests or tokens processed per unit time. Common units: QPS (queries per second), tokens/s. Inherent trade-off with latency.
TTFT (Time To First Token) From request arrival to the first output token — corresponds to the prefill phase. The most sensitive metric in interactive scenarios. See Latency, Throughput, and Concurrency.
TPOT (Time Per Output Token) Average generation time per token after the first — corresponds to the decode phase. The "typing speed."
E2E Latency (end-to-end latency) Total time from request to completed output. E2E = TTFT + TPOT × (output_tokens - 1).
QPS (queries per second) Requests processed per second — the request-level measure of throughput.
concurrency The number of requests served simultaneously. The bridge connecting latency and throughput. See Latency, Throughput, and Concurrency.
P50 / P95 / P99 (percentile latency) Latency percentiles: P50 is the median; P95 means 95% of requests finish below it; P99 captures the long tail. Production services are usually judged by P99, not the average.
Little's LawConcurrency = throughput × latency. Know any two, derive the third — the foundation of capacity planning for inference systems. See Latency, Throughput, and Concurrency.
SLA / SLO (service level agreement / objective) Service-level agreement/objective, usually stated as "P99 < X ms." The availability boundary of an inference system.
prefill / decode The two phases of LLM inference: prefill processes the entire prompt in one pass (compute-intensive); decode generates tokens autoregressively one at a time (memory-intensive). See Latency, Throughput, and Concurrency and The GPU Memory Hierarchy and the Bandwidth Wall.
time-to-first-token (TTFT) See TTFT.
inter-token latency (ITL) An alias for TPOT — the gap between consecutive tokens in streaming generation.
Memory & Compute
HBM (high bandwidth memory) GPU external memory: high bandwidth (A100 2 TB/s, H100 3.35 TB/s, H200 4.8 TB/s) and large capacity (80–192 GB), yet still far slower than on-chip SRAM. See The GPU Memory Hierarchy and the Bandwidth Wall and Hardware Primer.
SRAM (static random access memory) On-chip GPU shared memory — the L1/shared storage inside an SM. Bandwidth on the order of 10–30 TB/s but tiny capacity (A100: 228 KB shared memory per SM). The vehicle for tiling data inside a kernel.
L2 cache The GPU's on-chip level-2 cache, shared across all SMs. A100 L2 is 40 MB, H100 50 MB, B200 up to 60 MB. Hits avoid trips to HBM.
memory bandwidth The volume of data read from / written to HBM per unit time — the ceiling of the decode phase. The core bottleneck of LLM decode.
memory capacity The hard constraint on how large a model and how large a KV cache fit. A 70B model in FP16 needs ~140 GB — a single 80 GB card can't hold it, so tensor parallelism or quantization is required.
compute / FLOPS Floating-point operations per second. FP16: A100 312 TFLOPS, H100 989 TFLOPS, B200 2250 TFLOPS. See The Roofline Model and Compute Analysis and Hardware Primer.
Roofline model Unifies compute and bandwidth in one model: performance ceiling = min(peak compute, bandwidth × arithmetic intensity). See The Roofline Model and Compute Analysis.
arithmetic intensity Floating-point operations per byte of memory traffic, in FLOPS/Byte. Determines whether an operator is compute-bound or memory-bound. See The Roofline Model and Compute Analysis.
compute-bound The operator's arithmetic intensity is high and performance is capped by FLOPS. Typical: large-batch GEMM, the prefill phase.
memory-bound The operator's arithmetic intensity is low and performance is capped by bandwidth. Typical: small-batch GEMM, the decode phase, LayerNorm, activation functions. The defining characteristic of LLM decode.
capacity-bound Problems caused by memory not fitting: the model won't fit, the KV cache overflows, batch size can't grow. See The GPU Memory Hierarchy and the Bandwidth Wall.
roofline ridge The arithmetic intensity at the intersection of the compute line and the bandwidth line on a Roofline chart. Below the ridge → memory-bound; above → compute-bound.
MAC / FLOP (multiply-accumulate / floating-point operation) One multiply-accumulate = 2 FLOPs. The common unit for GEMM compute.
Quantization & Compression
PTQ (post-training quantization) Quantizing after training — no retraining required, fastest to adopt. AWQ, GPTQ, and SmoothQuant all fall in this family. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
QAT (quantization-aware training) Simulating quantization error during training — higher accuracy than PTQ, but needs training data and training budget. See Model Quantization Fundamentals.
INT8 / INT4 / FP8 / FP4 / FP16 / BF16 Common numeric precisions. FP16/BF16 are training defaults; INT8 is the inference mainstream; INT4 is the aggressive edge choice; FP8 gained native hardware support starting with Hopper; FP4 arrives with Blackwell. See Model Quantization Fundamentals and Hardware Primer.
weight-only quantization Compressing only the weights to INT4/INT8 while activations stay FP16. The mainstream LLM deployment scheme (GPTQ, AWQ), because activations are hard to quantize. See Weight-Only Quantization and Mixed Precision.
mixed precision Different layers / different tensors at different precisions. Common pattern: INT4 weights + FP16 activations + a few layers kept in FP16.
GPTQ A PTQ algorithm based on second-order Hessian information, quantizing weights column by column with error compensation. One of the de facto INT4 LLM quantization standards. See Weight-Only Quantization and Mixed Precision.
AWQ (activation-aware weight quantization) A PTQ method built on the observation that "some weight channels matter more": apply small protective rescales to important channels before quantizing. The other INT4 mainstream. See Weight-Only Quantization and Mixed Precision.
SmoothQuant Moves the "hard-to-quantize" part of activations to the weight side via smoothing rescales, enabling simultaneous INT8 for weights and activations. See Model Quantization Fundamentals.
ZeroQuant ByteDance's PTQ scheme combining per-token activation quantization with per-group weight quantization.
per-tensor / per-channel / per-group quantization granularity The sharing scope of quantization parameters (scale, zero-point). Per-tensor shares one set across the whole tensor (cheapest memory, biggest accuracy loss); per-channel gives each output channel its own; per-group gives each group of N consecutive elements its own (GPTQ/AWQ commonly use group=128).
symmetric / asymmetric quantization Symmetric quantization pins zero-point at 0, centering the range on zero (INT8 range [-127, 127]); asymmetric allows a zero-point offset to hug the distribution better (e.g., post-ReLU activations), at the cost of one extra zero-point correction per operation.
calibration The PTQ procedure of running a small representative dataset through forward passes to measure activation distributions and determine quantization parameters. See Model Quantization Fundamentals.
dequantize Restoring quantized weights to high precision for computation at inference time. In LLM deployment this usually happens on the fly inside the GEMM kernel (weight-only GEMM).
sparsification Zeroing out part of the weights, paired with sparse operators for acceleration. NVIDIA's 2:4 structured sparsity has native hardware support. See Pruning and Sparsification.
Operators & Graphs
operator (op) The basic node of a computation graph, corresponding to one explicit mathematical operation: MatMul, LayerNorm, Softmax, Conv. See Kernel Fusion and Custom Kernels.
kernel The concrete hardware implementation of an operator. The same operator deserves different kernels on different hardware / at different batch sizes — the starting point of kernel auto-tuning.
kernel fusion Merging multiple adjacent operators into a single kernel so intermediate results never fall back to HBM. See Kernel Fusion and Custom Kernels.
kernel auto-tuning Generating multiple candidate kernels for one operator and measuring them on the target hardware to pick the fastest. Built into TensorRT, TVM, and Triton. See Kernel Fusion and Custom Kernels.
FlashAttention Rewrites attention's softmax(QK^T)V as tiling + online softmax, avoiding materializing the N×N attention matrix. v1/v2/v3 are milestones of LLM acceleration. See Kernel Fusion and Custom Kernels and Classic Papers in Depth.
FlashInfer An attention kernel library purpose-built for LLM inference, covering prefill/decode/append scenarios; adopted by vLLM/SGLang.
Triton (OpenAI Triton) A Python DSL + compiler for writing GPU kernels that hides PTX details while achieving near-hand-written-CUDA performance. FlashAttention v2 onward is Triton-based. See Kernel Fusion and Custom Kernels.
Don't mix up the two Tritons
OpenAI Triton is a GPU kernel DSL/compiler for writing custom operators; NVIDIA Triton Inference Server is an inference serving framework for deploying models. Same name, completely unrelated things. The former is covered in Kernel Fusion and Custom Kernels; the latter in Triton Inference Server.
CUDA Graph Records a sequence of CUDA kernel calls into a graph launched with a single launch, saving CPU launch overhead. In LLM inference it's commonly used for fixed-shape decode.
operator fusion Same as kernel fusion.
IR (intermediate representation) The framework-agnostic computation-graph representation used inside compilers. ONNX, MLIR dialects, and Relax IR are all IRs.
ONNX (Open Neural Network Exchange) The cross-framework model exchange format — the de facto standard. PyTorch/TF/JAX models are exported into a unified ONNX graph, which engines then consume. See ONNX Runtime: Cross-Platform.
MLIR (multi-level intermediate representation) A multi-level IR framework under the LLVM project; PyTorch torch.compile and JAX/XLA are both built on it. The new trend.
XLA (accelerated linear algebra) Google's graph compiler, compiling HLO IR into GPU/TPU/CPU code. JAX and TF compile through XLA by default.
TVM Apache's end-to-end deep learning compiler, covering graph optimization + kernel auto-tuning + cross-hardware backends. See Computation Graph Optimization.
Relax The next-generation LLM inference IR from the TVM team, purpose-designed for dynamic shapes and large models.
torch.compile (PyTorch 2.x) PyTorch 2.0's JIT compilation entry point, built on TorchDynamo + Inductor underneath — capable of graph capture, fusion, and kernel generation.
layout (data layout) How tensors are arranged in memory: NCHW / NHWC / blocked. Different hardware prefers different layouts; TensorRT's layout conversions are a major optimization focus.
Systems & Serving
KV cache (key-value cache) Caching the K and V of historical tokens during autoregressive generation, so every new token doesn't recompute the entire history. The biggest memory consumer in LLM inference. See vLLM and PagedAttention and The GPU Memory Hierarchy and the Bandwidth Wall.
PagedAttention vLLM's KV cache management scheme, borrowing OS virtual memory: KV split into fixed-size blocks, allocated on demand, with prefix sharing. See vLLM and PagedAttention.
continuous batching (continuous / dynamic batching) Requests join a batch as they arrive and finished ones leave at any time — no waiting for the batch to align. The key to LLM serving throughput. See Batching and Request Scheduling and vLLM and PagedAttention.
static batching Requests accumulate into a batch, enter together, and leave together. Simple but low-throughput, badly dragged by long tails.
prefix caching Reusing prefill KV for requests that share the same system prompt. Supported by vLLM, SGLang, and TGI; significantly cuts TTFT in multi-tenant scenarios. See Batching and Request Scheduling.
speculative decoding A small model (the draft) guesses several tokens; the large model verifies them in parallel, and hits yield multiple tokens per pass. Medusa and EAGLE-2/3 are the representatives. See Speculative Decoding and Medusa/EAGLE.
tensor parallelism (TP) Sharding each layer's weights into N parts across N GPUs; each computes a slice and AllReduce aggregates. Communication-intensive; requires NVLink. See Distributed Inference (TP/PP).
pipeline parallelism (PP) Sharding the model by layers into N stages across N GPUs; requests flow through like an assembly line. Low communication but prone to bubbles; usually paired with micro-batching. See Distributed Inference (TP/PP).
expert parallelism (EP) Parallelism specific to MoE models: different experts live on different GPUs. Core to DeepSeek-V3 and Mixtral deployments.
AllReduce / AllGather / All-to-All Collective communication primitives. AllReduce reduces then broadcasts (used by TP); AllGather collects pieces from every rank (used by PP); All-to-All is a full exchange (used by EP). See Distributed Inference (TP/PP).
vLLM UC Berkeley's open-source LLM inference engine — started from PagedAttention + continuous batching and has the largest community. See vLLM and PagedAttention.
TensorRT-LLM NVIDIA's official LLM inference engine — the LLM-specialized branch of TensorRT, with FP8, in-flight batching, and speculative decoding all built in. See TensorRT-LLM.
SGLang An LLM inference framework started at Berkeley/MSU, focused on structured generation and RadixAttention prefix caching. See Inference Engine Comparison.
TGI (Text Generation Inference) Hugging Face's LLM serving engine, deeply integrated with the Transformers ecosystem.
Triton Inference Server NVIDIA's multi-backend inference serving framework (not OpenAI Triton) — unified management of TensorRT/PyTorch/TF models with HTTP/gRPC and dynamic batching. See Triton Inference Server.
llama.cpp A C++ implementation of LLM inference that runs on CPU/GPU across all platforms; its GGUF format is the de facto on-device deployment standard. See llama.cpp and GGUF.
ONNX Runtime Microsoft's cross-platform inference engine covering CPU/GPU/NPU/mobile. See ONNX Runtime: Cross-Platform.
OpenVINO Intel's inference optimization toolchain for CPU/iGPU/VPU. See OpenVINO and CPU Inference.
engine (inference engine artifact) The executable TensorRT produces after compiling a model (.engine/.plan), bound to a specific GPU architecture and precision. See TensorRT and GPU Inference.
plan file Same as the engine — TensorRT terminology.
in-flight batching TensorRT-LLM's continuous-batching implementation — same idea as vLLM's continuous batching.
chunked prefill Splitting the prefill of a long prompt into multiple chunks scheduled together with decode, so a long prompt doesn't monopolize the GPU. Proposed by the SARATHI paper; adopted by both vLLM and TensorRT-LLM.
Hardware
GPU (graphics processing unit) A massively parallel compute chip — the workhorse of AI training and inference. See GPU Architecture and Optimization and Hardware Primer.
SM (streaming multiprocessor) The basic compute unit of a GPU. Each SM contains multiple CUDA cores, Tensor cores, shared memory, and schedulers. A100: 108 SMs; H100: 132; B200: 148.
CUDA core The GPU's scalar compute unit — FP32/FP64 math runs on it. A100 has 64 FP32 CUDA cores per SM.
Tensor Core Matrix multiply-accumulate units introduced with NVIDIA Volta, completing one 4×4×4 matmul per clock. All FP16/BF16/INT8/FP8/FP4 throughput comes from these. See GPU Architecture and Optimization.
warp The smallest scheduling unit of an SM — 32 threads executing the same instruction in lockstep (SIMT). Warp divergence degrades efficiency.
SIMT (single instruction, multiple threads) The GPU execution model: one instruction, multiple threads — between SIMD and SMT.
NPU (neural processing unit) A generic term for chips dedicated to neural-network inference — broadly includes TPUs, Ascend, Apple Neural Engine, Qualcomm Hexagon, and more.
TPU (tensor processing unit) Google's custom AI ASICs — TPU v4/v5e/v5p. Natively supported by JAX/XLA.
Ascend 910B Huawei's in-house AI processor with the Da Vinci architecture — the mainstay of domestic Chinese compute.
Ampere / Hopper / Blackwell Three generations of NVIDIA datacenter GPU microarchitectures: A100 (2020), H100 (2022), B200 (2024). See Hardware Primer.
A100 / H100 / H200 / B200 / GB200 Specific NVIDIA datacenter GPU models. See the comparison table in Hardware Primer.
cuDNN NVIDIA's deep-learning operator library — the foundational implementations of convolutions, pooling, normalization, and more.
cuBLAS NVIDIA's BLAS matrix-multiplication library — the bottom layer of GEMM.
NCCL (NVIDIA collective communications library) NVIDIA's multi-GPU collective communication library — high-performance AllReduce/AllGather. Underlies PyTorch DDP/FSDP and TP/PP.
NVLink / NVSwitch High-speed GPU interconnects: NVLink 4.0 at 100 GB/s per direction (H100); NVSwitch is the all-to-all switch chip — on B200, the NVLink Switch. See Distributed Inference (TP/PP) and Hardware Primer.
PCIe The general-purpose bus for CPU-GPU and multi-GPU interconnect — PCIe Gen5 ×16 at 64 GB/s per direction. Slower than NVLink but universal.
TDP (thermal design power) The chip's thermal design power. A100 400W, H100 SXM 700W, B200 1000W+ — determines rack and cooling design.
MIG (multi-instance GPU) Hardware-level partitioning of one GPU into isolated instances since Hopper; supported on A100/H100 for multi-tenant scenarios.
DVFS (dynamic voltage and frequency scaling) Dynamic voltage/frequency scaling — saves power but jitters latency, often disabled for latency-sensitive inference.
Easily Confused — Side by Side
Inference / Training / Fine-Tuning
Training learns parameters from data; fine-tuning continues training on existing weights to change parameters; inference uses parameters for prediction without changing them. The three share PyTorch, but their deployment shapes differ completely — training uses the framework's eager mode, inference uses dedicated engines (TensorRT/vLLM). See Inference vs. Training vs. Fine-Tuning.
Latency / Throughput / Concurrency
Latency is how fast a single request is; throughput is the total volume per unit time; concurrency is how many can be served at once. Little's Law binds them: concurrency = throughput × latency. Before optimizing, ask "does the business want low latency or high throughput" — single-user chat cares about TTFT; offline batch processing cares about tokens/s. See Latency, Throughput, and Concurrency.
compute-bound / memory-bound / capacity-bound
Compute-bound: not enough FLOPS — add compute or cut operations; memory-bound: not enough bandwidth — add bandwidth or cut traffic; capacity-bound: memory doesn't fit — add memory or quantize. LLM inference: prefill is usually compute-bound, decode is usually memory-bound, long context is often capacity-bound. The Roofline model is the judging tool. See The Roofline Model and Compute Analysis.
PTQ / QAT / weight quantization / activation quantization
PTQ quantizes after training — fastest to adopt; QAT simulates quantization during training — highest accuracy but expensive. Weight quantization compresses weights only and keeps activations high-precision — the LLM deployment mainstream; activation quantization also compresses activations and needs migration techniques like SmoothQuant. LLM activations are full of outliers, so pure activation quantization loses significant accuracy — the fundamental reason weight-only became the mainstream. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
vLLM / TensorRT-LLM / SGLang / TGI
All four are LLM inference engines with different positions: vLLM — largest open-source community, started from PagedAttention; TensorRT-LLM — NVIDIA official, deeply fused with FP8 hardware; SGLang — focused on structured generation and RadixAttention prefix caching; TGI — from Hugging Face, closest to the Transformers ecosystem. Selection trade-offs in Inference Engine Comparison and Benchmark Data & Tool Profiles.
The two Tritons
OpenAI Triton is a GPU kernel DSL (for writing custom operators — see Kernel Fusion and Custom Kernels); NVIDIA Triton Inference Server is an inference serving framework (for deploying models — see Triton Inference Server). Same name, different things — context is the only way to tell them apart in the wild.
Further Reading
- What Is Inference Acceleration? — the starting point of the site's concepts
- Anatomy of the Overall Architecture — where each term sits in the system
- Hardware Primer — concrete hardware parameters behind the terms
- Benchmark Data & Tool Profiles — measured data for engines and hardware
- Curated Resources — materials for going deeper
- Common Pitfalls and Anti-Patterns — engineering lessons behind the terms