Skip to content

Glossary

At a glance Core terminology for model deployment: fundamentals, models and formats, compression and optimization, serving and architecture, hardware and performance, monitoring and operations, and LLM inference—precise definitions and distinctions for 60+ terms.

Glossary ​

How to Use This Glossary

This glossary is grouped by topic, and you don't need to read it front to back—treat it like a dictionary: when you hit an unfamiliar term, look it up directly with browser search (Ctrl+F). The end of each section untangles easily confused terms; entries with a "See also" link take you to the corresponding chapter for a deeper read.

Fundamentals ​

Model Deployment The end-to-end process of putting a trained model into production in a form that is servable, scalable, and observable—covering model conversion and optimization, service wrapping, release, monitoring, and iteration. Deployment is not a single "export the model file" step; it is an ongoing lifecycle. See What Is Model Deployment.

Inference The process of using a trained model to make predictions on new input data. Inference and training share the same forward pass, but their constraints are completely different: inference cares about latency, throughput, and cost, and typically runs low-precision, compiler-optimized models.

Forward Pass The computation that carries data layer by layer from the input layer to the output layer. Training uses it to compute the loss; inference keeps only the forward pass (no backpropagation, no gradients).

Model Serving The engineering practice of wrapping a model into an online service accessible over HTTP/gRPC or similar interfaces, covering request/response protocols, batching, model loading and lifecycle management, health checks, and scaling. See Serving and Inference APIs.

Inference Engine The software runtime that actually executes model computation, such as ONNX Runtime, TensorRT, OpenVINO, and llama.cpp. It maps the computation graph onto low-level operator libraries (cuDNN, oneDNN, etc.) and hardware—the execution layer beneath "model serving".

MLOps The engineering culture and methodology of applying DevOps practices (CI/CD, automation, monitoring, observability) to the full machine learning lifecycle, covering data, training, evaluation, deployment, monitoring, and governance. Deployment is one step of MLOps, not all of it. See MLOps Deployment Pipelines.

Model Registry A component that centrally manages model artifacts (versions, metadata, lineage, and status such as staging/production). It is the "gate" a model passes through from experiment to production. Representative implementations: MLflow Model Registry, Hugging Face Hub, Kubeflow Hub.

Data Drift A change in the distribution of input data relative to the training set—for example, a shift in the user population or feature value ranges. Data drift is a common early warning sign of model performance degradation.

Concept Drift A change in the mapping between inputs and outputs—"same input, different correct answer". For example, after risk-control rules tighten, the default probability of users with identical features rises sharply.

Model Drift The degradation of a model's actual prediction performance over time, the combined effect of data drift, concept drift, and other factors. In production it is usually detected indirectly through metrics such as PSI and KS.

Inference vs. inference engine Inference is the act/process of model computation (a single forward pass); an inference engine is the software that performs it. Analogy: inference is the act of "running"; the inference engine is the "legs". Serving frameworks (Triton, FastAPI, etc.) handle scheduling and protocols, while the inference engine does the math.

Data drift vs. concept drift One-line memory aid: data drift means "the inputs changed"; concept drift means "the rules changed". Both can surface as anomalies in production metrics (PSI, KS, etc.), but the remedies differ—the former calls for resampling or more data, the latter usually for retraining.

Models and Formats ​

ONNX (Open Neural Network Exchange) An open, vendor-neutral model exchange format that describes model structure as a computation graph, enabling interoperability between frameworks like PyTorch and TensorFlow and inference engines. See Model Formats and Conversion.

TorchScript PyTorch's static-graph serialization format (.pt/.torchscript). Generated via torch.jit.trace or script, it can run without Python. Compared to ONNX it covers PyTorch operators more completely, but its ecosystem interoperability falls short.

TFLite (TensorFlow Lite) TensorFlow's lightweight format for mobile and embedded devices (.tflite), supporting INT8/FP16 quantization and running on Android, iOS, and MCUs via TFLite Runtime / LiteRT.

TensorRT engine (.engine / .plan) A highly specialized engine file produced by compiling an ONNX or similar model with NVIDIA TensorRT. It is bound to a specific GPU architecture, batch size, and precision configuration—switching GPUs or changing settings forces a rebuild. See TensorRT and Edge Deployment.

GGUF The model weight format of the llama.cpp ecosystem, packaging weights, tokenizer, and hyperparameters into a single file with support for 2–8 bit integer quantization, making it easy to run local LLMs on CPUs and low-VRAM devices.

SafeTensors A weight storage format championed by Hugging Face. Its key property is that the file header records each tensor's offset and shape, so loading never requires reading the whole file into memory. It is safer than PyTorch's pickle (no arbitrary code execution) and loads faster.

opset (Operator Set) The version number of the ONNX operator set. Each version adds or modifies operator definitions; opset=17 means the model uses ONNX operator set 17. During conversion, an opset that is too high may exceed what the target runtime supports.

Dynamic Axes Dimensions that are allowed to vary at inference time, typically batch size and sequence length. Dynamic axes add flexibility but reduce the optimization headroom of some engines (TensorRT, for example).

Computational Graph A representation of the model as a directed acyclic graph of "tensors flowing through operators"—the common abstraction behind ONNX, TorchScript, TensorRT, and similar formats. Graph optimizations (operator fusion, constant folding) all happen at this layer.

Intermediate Representation (IR) A model representation sitting between the "source framework graph" and "target hardware instructions". ONNX is a general-purpose IR; OpenVINO IR (.xml+.bin) is a specialized IR for its runtime. IR is a core concept in the layered design of compilers and inference engines.

Compression and Optimization ​

Quantization Approximating weights and/or activations at lower precision (INT8, FP16, INT4, etc.) in exchange for a smaller model footprint, lower memory bandwidth, and faster computation. Accuracy loss can be controlled through calibration and algorithmic improvements. See Quantization.

PTQ (Post-Training Quantization) Quantizing directly after using a small calibration dataset to profile activation distributions—no retraining required. Low cost and quick to ship, it is the first choice for the vast majority of deployment scenarios; escalate to QAT when accuracy is critical.

QAT (Quantization-Aware Training) Simulating quantization error during training/fine-tuning (inserting fake-quantize operators on weights) so the model learns to tolerate low precision. The best accuracy, but it requires training resources.

FP16 16-bit floating point (1 sign bit + 5 exponent bits + 10 mantissa bits). Halves memory and bandwidth compared to FP32; the default precision for GPU inference.

BF16 (Brain Float 16) A 16-bit float that keeps FP32's 8-bit exponent and truncates only the mantissa. Its dynamic range matches FP32, making it suitable for training and numerically sensitive inference; the tradeoff is lower mantissa precision. Supported mainly on Ampere and newer architectures.

INT8 8-bit integer quantization. Weights and activations are computed in INT8, usually with per-channel/per-tensor scale factors (scale/zero-point). Significantly better speed and energy efficiency than FP16, but it requires calibration.

GPTQ (GPT Quantization) A post-training weight-only quantization method for LLMs (second-order error minimization from a single calibration sample), typically producing 4-bit weights (W4A16). Introduced in the paper GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. See Classic Quantization Papers.

AWQ (Activation-aware Weight Quantization) A post-training weight quantization method for LLMs, built on the idea of "protecting the few critical weight channels that most affect activations". At 4 bits it beats GPTQ on quality and is more robust to the choice of calibration set.

Knowledge Distillation Training a small model (student) to mimic the output distribution of a large model or ensemble (teacher)—the classic way to shrink a model while preserving capability. See Distillation, Pruning, and Low-Rank Factorization.

Pruning Removing unimportant weights or neurons (structured/unstructured) to cut computation and model size. Unstructured pruning needs sparse operator support before the speedup actually materializes.

Sparsity Zeroing out large portions of the weight matrix and exploiting sparse formats (2:4 sparsity, etc.) to skip computations on zeros. NVIDIA delivers roughly 2x acceleration through structured sparsity on Ampere architectures.

Low-Rank Factorization Decomposing a large weight matrix into the product of two smaller matrices (e.g., via SVD) to reduce parameters and computation. LoRA's "low-rank adaptation" idea on the training side shares the same origin.

PTQ vs. QAT One-line decision rule: try PTQ first—it needs no training and gives you a result in five minutes. If accuracy falls short and resources allow, move up to QAT. QAT has a higher accuracy ceiling but costs an order of magnitude more.

Serving and Architecture ​

Online Inference The real-time request/response service model: clients send a request and the server returns a prediction synchronously, demanding low latency (usually p99 under a few hundred milliseconds). Typically implemented as HTTP/gRPC microservices. See Deployment Architecture Patterns.

Batch Inference (Offline) Feeding large volumes of samples through the model in batched passes to maximize throughput and amortize overhead. Suited to offline labeling and ETL-style tasks where latency does not matter.

Dynamic Batching The server automatically accumulates requests that arrive within a short window and executes them together once a batch fills up—raising GPU utilization while keeping the service online. The throughput comes at the cost of a maximum delay budget (max_batch_delay).

Continuous Batching Scheduling for streaming LLM generation: requests enter and leave the batch at token granularity, and a finished request immediately yields its slot to a new one instead of waiting for the whole batch to finish in lockstep. It is the core mechanism behind the LLM throughput gains of vLLM, TGI, and similar engines. See LLM Inference Optimization.

Streaming Inference The server pushes partial results (tokens/chunks) to the client as they are generated (SSE / gRPC stream), dramatically cutting the user's perceived time to first token. The de facto standard for LLM chat.

Edge Deployment Running models on devices close to the data source (phones, cameras, in-vehicle units, Jetson, etc.) in exchange for low latency, offline availability, and data privacy. See TensorRT and Edge Deployment.

Serverless A platform model with per-request billing and automatic scale-to-zero (AWS Lambda, KServe Serverless mode, etc.). Suited to highly variable traffic without extreme latency requirements; cold starts are the main price.

Load Balancing Distributing requests across multiple model replicas to even out load and tolerate single-point failures. In inference scenarios it is often combined with routing strategies that are aware of queue depth/memory usage.

Canary Release Exposing the new version to a small slice of traffic first, then ramping up gradually once nothing looks wrong. The safest way to ship a model to production.

Blue-Green Deployment Running old and new environments side by side and switching entry traffic for second-level releases and rollbacks. Two sets of resources coexist, so it costs more, but rollback is the simplest.

A/B Testing Splitting traffic between different models (versions) by policy and using business metrics (conversion rate, satisfaction, etc.) to statistically evaluate which one performs better. The difference from canary release: canary is about "shipping safely", A/B is about "comparing effectiveness".

Batch inference vs. dynamic batching Offline batch inference means "the data side accumulates a batch before computing"—users simply wait for results. Dynamic batching means "the serving side batches concurrent requests"—the goal is squeezing the GPU dry at the price of a little latency. The former is task-oriented; the latter handles real-time traffic.

Online inference vs. Serverless Online inference is an architecture (always-on service, low latency); Serverless is a platform billing/scaling model (per-invocation billing, scale-to-zero). You can host online inference on a Serverless platform, but Serverless is also often used for batch processing or event-driven entry points.

Hardware and Performance ​

Memory Bandwidth How much data GPU memory can read per second (GB/s). For memory-bound workloads like LLM decoding, memory bandwidth directly caps token generation speed. See GPUs and Hardware Selection.

TFLOPS (Tera FLOPs per Second) Trillions of floating-point operations per second—a measure of a GPU's peak compute. Compute-bound operators like convolutions and matrix multiplies are bottlenecked by compute (compute-bound).

KV Cache The intermediate tensors cached during LLM inference to avoid recomputing K/V attention for past tokens. It grows linearly with sequence length, is the main driver of LLM memory usage, and is the target of memory optimizations like PagedAttention.

TTFT (Time To First Token) Time from sending a request to receiving the first output token. Dominated by the prefill phase, it is the key metric for "how fast does it respond".

TPOT (Time Per Output Token) The average time to generate each output token. Decoding is token-by-token and serial, so TPOT determines how smooth the typewriter effect feels. Total output latency ≈ TTFT + TPOT × (output length − 1).

P99 (99th Percentile) The latency value at the 99th percentile when sorted, representing the experience of the "worst 1% of users". Service performance is usually described with the P50/P99 pair; P99 is more sensitive to the long tail.

QPS (Queries Per Second) Requests processed per second—the intuitive measure of service capacity. Load testing and capacity planning revolve around "does P99 hold at the target QPS". See Load Testing and Capacity Planning.

Throughput The total number of samples/tokens processed per unit time (tokens/s, requests/s). Throughput and latency often work against each other: dynamic batching raises throughput but also latency, and load testing is where you find the balance. See Performance Optimization and Capacity Planning.

NVLink NVIDIA's high-speed inter-GPU interconnect, with bandwidth far exceeding PCIe. Communication-heavy distributed inference such as tensor parallelism relies on NVLink for near-linear scaling.

CUDA NVIDIA's GPU computing platform and programming model. Matching CUDA versions, driver versions, and cuDNN/cuBLAS libraries is one of the top deployment troubleshooting headaches (see the GPU environment section of the Systems Fundamentals Cheat Sheet).

Monitoring and Operations ​

PSI (Population Stability Index) A metric measuring the overall shift between two distributions, commonly used to monitor feature/score distribution drift. Rule-of-thumb thresholds: <0.1 stable, 0.1–0.25 worth watching, >0.25 severe drift. See Monitoring and Observability.

KS Test (Kolmogorov–Smirnov Test) A non-parametric hypothesis test comparing the maximum cumulative difference between two sample distributions. In model monitoring it is used to judge whether a distribution has drifted statistically significantly, and also to measure how well a score separates positive from negative samples.

SLI (Service Level Indicator) A quantified measure of service quality, such as 99.9% availability, P99 latency, error rate, QPS. It is the measurement basis for SLOs.

SLO (Service Level Objective) The service quality target a team commits to internally, such as "P99 latency < 500ms with monthly attainment ≥ 99.5%". SLOs define "what counts as good" and drive alerting and capacity planning.

SLA (Service Level Agreement) Contractual service terms signed with external customers, usually looser than internal SLOs (leaving yourself margin). Breaches carry compensation/penalties.

Observability The ability to answer "why did the system end up in this state" through the three pillars (three signals) of metrics, logs, and traces. Model serving adds model quality metrics on top (accuracy, drift, LLM quality scores). See Putting Observability into Practice.

OpenTelemetry (OTel) The CNCF-incubated open observability standard, unifying the APIs, SDKs, and collection protocol (OTLP) for Trace/Metric/Log. It is the common layer for integrating with backends such as Prometheus, Grafana, and Jaeger.

Prometheus The CNCF time-series database and monitoring system. It scrapes metrics via a pull model, queries with PromQL, and alerts through Alertmanager—the de facto standard for Kubernetes and inference service monitoring.

Alerting The mechanism of triggering notifications based on rules (thresholds, SLO burn rates, etc.). Paired with SLOs, alerts should be actionable: receive an alert → know what to do. Avoid alert fatigue from floods of irrelevant notifications.

SLO vs. SLA An SLO is the internal engineering goal you set for yourself (you can tighten it and iterate); an SLA is a contract promised externally (with buffer and consequences). Common practice: keep the SLA slightly looser than the SLO, so external commitments are not shackled to internal targets.

LLM Inference ​

Autoregressive Generation How LLMs predict output token by token: each new token depends on all previous tokens. Decoding is serial per token and must carry the entire history—this is the fundamental reason KV Cache and continuous batching exist.

Speculative Decoding A small "draft model" guesses several tokens at once, then the large model verifies them in parallel—accepting multiple steps on success, rolling back on failure. While keeping the output distribution equivalent, it can speed up decoding by 2–3x.

Prefix / Prompt Caching Caching the KV of previously computed common prefixes (system prompt, conversation history) so new requests reuse them directly, significantly cutting TTFT and compute. vLLM's Automatic Prefix Caching (APC) is exactly this mechanism.

Tensor Parallelism Splitting a single layer's weights along dimensions across multiple GPUs, which then compute the same layer collaboratively. Communication is frequent and depends on NVLink/high-speed networking; it is the default parallelism for single-node multi-GPU inference.

Pipeline Parallelism Cutting the model into segments by layer across different GPUs, with data flowing through each in turn (micro-batch pipelining). Suited to deploying very large models across machines; communication is far lower than tensor parallelism, but pipeline bubbles exist.

PagedAttention The KV Cache paging mechanism proposed by vLLM: allocate KV in fixed-size blocks, physically non-contiguous and on demand, eliminating memory fragmentation and supporting cross-request sharing of common prefixes. It is the foundation vLLM is built on. See the PagedAttention Paper Deep Dive.

Prefill / Decode The two phases of LLM generation: Prefill processes the entire prompt in one parallel pass, producing the first token and filling the KV Cache; Decode generates token by token while updating the KV. The two phases differ sharply in compute and memory-access characteristics, forming the basis of disaggregated serving and scheduling optimizations.

Sampling Parameters (Top-k / Top-p / Temperature) Parameters controlling decoding randomness: Top-k keeps only the k highest-probability tokens, Top-p keeps only tokens whose cumulative probability reaches a threshold, and Temperature scales how peaked the probability distribution is. They are universal parameters of inference APIs and do not affect memory or throughput (only output quality).

LoRA (Low-Rank Adaptation) Low-rank adaptation: freeze the original weights and train only two small low-rank matrices as the delta. At deployment, the delta can be merged into the base model or loaded dynamically, enabling low-cost switching between multiple models.

Context Window The maximum number of tokens a model can process at once (input + output). Long contexts depend on more efficient attention and larger KV Cache management—one of the main threads of long-text inference optimization.

KV Cache and memory During LLM inference, memory must hold not just the model weights but also a KV Cache for every generated token of every concurrent request. That is why "a mere 7B model needs 40GB of VRAM"—the higher the concurrency and the longer the context, the bigger the KV. Quantizing the KV Cache and PagedAttention both squeeze space out of this "dynamic memory".

Further Reading ​