Appearance
Anatomy of the Overall Architecture
What exactly makes up an "inference system"? Most beginners picture "load the model, run the forward pass," but a real production-grade LLM service is far more than that. This page dissects a complete inference system into six layers, explains each in one sentence — what it does, why it exists, and where the pitfalls are — and links to the corresponding pages across the site. This page is the map of the whole site — after reading it, you own the skeleton of the field.
1. The Big Picture
text
┌─────────────────┬───────────────────────────────────────────────────────────────────────────────────────┐
│ ⑥ Application │ API gateway · streaming · Function Calling · RAG · Agent │
│ │ the boundary between user and model: turning token streams into product experience │
├─────────────────┼───────────────────────────────────────────────────────────────────────────────────────┤
│ ⑤ Orchestration │ request scheduling · batching · routing · multi-model routing · caching │
│ │ the scheduler of requests and models: keep the GPU fed, send each request its own way │
├─────────────────┼───────────────────────────────────────────────────────────────────────────────────────┤
│ ④ Engine │ vLLM · TensorRT-LLM · SGLang · TGI · Triton · llama.cpp │
│ │ turn model weights into a callable service: run forward, manage the KV cache │
├─────────────────┼───────────────────────────────────────────────────────────────────────────────────────┤
│ ③ Operator │ FlashAttention · FlashInfer · kernel fusion · custom kernels │
│ │ extreme optimization of a single operator: close the memory-vs-compute gap │
├─────────────────┼───────────────────────────────────────────────────────────────────────────────────────┤
│ ② Model │ quantization · pruning · distillation · sparsification │
│ │ make the model itself smaller: less memory, less traffic, higher compute density │
├─────────────────┼───────────────────────────────────────────────────────────────────────────────────────┤
│ ① Hardware │ GPU (H100/A100/L40) · NPU (Ascend/ANE) · CPU (AMX/AVX) │
│ │ physical compute and bandwidth: the ceiling of every optimization │
└─────────────────┴───────────────────────────────────────────────────────────────────────────────────────┘
▲ the higher the layer, the closer to the user and the more engineering; the lower the layer, the closer to hardware and the more algorithmic ▼The six layers are not isolated — upper layers depend on the capabilities of lower layers, and lower layers realize their value through upper layers. A real production LLM service (such as the ChatGPT or Claude APIs) is usually the result of all six layers optimized together. Let us dissect them one by one.
2. The Six Layers
Layer 1 Hardware: Physical Compute and Bandwidth
The ceiling of all optimization. Inference hardware falls into three major categories:
- GPU: the de facto standard for LLM inference. NVIDIA H100 (80 GB HBM3, 2 TB/s bandwidth, FP8 Tensor Cores), A100 (80 GB HBM2e, 1.5 TB/s), L40/L40S (48 GB, cost-effective cards), and the consumer RTX 4090 (24 GB, common for community deployments). See GPU Architecture and Optimization and Hardware Primer.
- NPU: alternatives for domestic and on-device scenarios. Huawei Ascend 910B, Apple Neural Engine (M-series chips), Cambricon MLU, and others — inference-friendly but weakly supported for training.
- CPU: edge deployment and extreme low-cost scenarios. Intel AMX (from Sapphire Rapids), AVX-512/VNNI, paired with OpenVINO and llama.cpp, can run a 7B model on a server CPU.
Why hardware is the ceiling of optimization
Every upper-layer optimization is bounded by the hardware's physical parameters. The two most important hardware metrics for LLM inference: HBM bandwidth (sets the token/s ceiling of the decode phase) and memory capacity (decides how large a model and how much KV cache fit). An A100 80 GB with 1.5 TB/s bandwidth caps a single-request Llama-2-70B FP16 at about 1.5 TB/s ÷ 140 GB ≈ 10.7 tokens/s — this is why single-request LLM inference cannot get "fast" by optimizing compute alone; it can only reduce memory (quantization) or raise concurrency (batching). See The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis.
Layer 2 Model: Make the Model Smaller
Model-layer optimization does not change the inference engine — it changes the model itself. Three main techniques:
- Quantization: FP16 → INT8/INT4/FP8, cutting memory, memory traffic, and compute all at once — the highest-ROI optimization. For mainstream LLM quantization schemes see Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
- Pruning: remove unimportant weights/channels/heads. Structured pruning (removing whole rows and columns) genuinely saves compute; unstructured pruning needs sparse-kernel support and is hard to productionize on LLMs. See Pruning and Sparsification.
- Distillation: use a large model to teach a small one, letting the small model approach the large model's quality. Distillation is costly for LLMs but remains important for classification/retrieval tasks. See Knowledge Distillation.
The model layer is the precondition for whether the upper-layer engine can run at all — without quantizing a 70B model to INT4, it does not fit on a single 80 GB A100; and if it does not fit, you need tensor parallelism, which immediately raises complexity by an order of magnitude.
Layer 3 Operator: Make a Single Operator Fast
The operator layer — "write optimized kernels for the single hottest operator" — is the technical deep water of inference engineering.
- FlashAttention-2/3: compute attention's intermediate QK^T matrix in SRAM tiles instead of shuttling through HBM — a 2-4x single-operator speedup and absolute standard equipment for LLM inference. See Kernel Fusion and Custom Kernels.
- FlashInfer: an operator library from UC Berkeley that unifies prefill/decode attention, widely adopted by vLLM/SGLang/TensorRT-LLM.
- Handwritten CUDA/Triton kernels: one of vLLM and SGLang's core competitive edges is their custom attention/rmsnorm/rope kernels.
The key judgment at the operator layer: when GPU utilization is <30%, the bottleneck is the operator and you need handwritten kernels; when GPU utilization is >60%, the bottleneck is scheduling and you need batching. This judgment decides which layer you should optimize.
Layer 4 Engine: Turn Weights into a Service
The engine layer is where inference engineers spend most of their day — it packages "model weights + hardware + operators + scheduling" into a callable service.
| Engine | Home turf | Strengths | Weaknesses |
|---|---|---|---|
| vLLM | Open-source LLM serving | PagedAttention, continuous batching, easy to use | Operators tuned less aggressively than TensorRT-LLM |
| TensorRT-LLM | NVIDIA's closed-source extreme | Fastest on H100, best FP8/FP4 support | Complex to deploy, small community |
| SGLang | Structured generation, agents | RadixAttention (prefix sharing), compressed FSM | Smaller edge in general chat |
| TGI | HuggingFace ecosystem | Seamless with the HF model hub | Slightly behind vLLM in performance |
| Triton Inference Server | Multi-model, multi-framework | Multi-model management, enterprise-grade | Less LLM-tuned than specialized engines |
| llama.cpp | CPU / edge | Extreme cross-platform support, great quantization | Insufficient for high-traffic scenarios |
| ONNX Runtime | Non-LLM models | Cross-hardware, mature ecosystem | Weak LLM support |
| OpenVINO | Intel hardware | Great CPU/iGPU optimization | Intel ecosystem only |
The core logic of engine selection is in Inference Engine Comparison. One-sentence summary: NVIDIA GPU + high traffic → vLLM/TensorRT-LLM; CPU/edge → llama.cpp/OpenVINO; mixed multi-model → Triton Inference Server.
Layer 5 Orchestration: Feed Many Requests at Once
The orchestration layer is the core innovation of LLM inference engineering and the layer inference engineers care most about after 2023.
- Continuous batching: let requests of different lengths join and leave the batch dynamically, lifting GPU utilization from <5% (single-request decode) to 60%+. See Batching and Request Scheduling.
- KV cache management: PagedAttention pages the KV cache to eliminate fragmentation, raising concurrency from single digits to dozens. See the vLLM case study.
- Request routing: simple requests go to a small model, complex ones to a large model, cutting overall cost by 50%+. See Model Serving and Orchestration.
- Speculative decoding: a draft head guesses a few tokens and the large model verifies them in parallel — 2-6x speedup without changing the output distribution. See Speculative Decoding and Medusa/EAGLE.
- Multi-model routing: A/B testing, canary releases, traffic splitting by ratio.
The value of the orchestration layer: raise throughput 10-100x without changing model weights or swapping hardware. This is the core competitive edge of vLLM and SGLang.
Layer 6 Application: Turn Token Streams into a Product
The application layer is the boundary between user and model, and it determines the product shape of the inference system:
- API gateway: authentication, rate limiting, billing, logging — wrapping the inference engine into a stable API.
- Streaming output: SSE/WebSocket lets users watch tokens appear one by one; time-to-first-token (TTFT) matters more than last-token latency. See Latency, Throughput, and Concurrency.
- Function Calling / Tool Use: the model emits structured calls, the application layer executes them and writes back — requiring the engine to support structured-output acceleration (such as SGLang's compressed FSM).
- RAG (retrieval-augmented generation): retrieve → assemble prompt → infer; the prompt prefix can be shared and cached via RadixAttention, see vLLM/SGLang.
- Agent: multi-round inference + tool calls + loops, placing new demands on the engine's latency and concurrency.
The application layer determines the direction of optimization: streaming scenarios prioritize TTFT; batch scenarios prioritize total throughput; agent scenarios prioritize prefix sharing and concurrency. Quoting "which engine is fastest" without the application shape is meaningless.
3. How the Six Layers Map to Roles
Different roles distribute their energy very differently across the six layers, mapping directly to the job landscape in Module Guide and Job Landscape:
| Layer | Inference Engineer | Algorithm Engineer | Infrastructure Engineer | Application Engineer |
|---|---|---|---|---|
| ⑥ Application | ★★ | ★ | ★★ | ★★★★★ |
| ⑤ Orchestration | ★★★★★ | ★★ | ★★★★ | ★★ |
| ④ Engine | ★★★★★ | ★★★ | ★★★★ | ★ |
| ③ Operator | ★★★★ | ★★★★ | ★★ | — |
| ② Model | ★★★★ | ★★★★★ | ★ | ★ |
| ① Hardware | ★★★ | ★★★ | ★★★★★ | — |
Advice for beginners: layers ⑤ and ④ are the "main battlefield" of the inference engineer — ⑤ decides the throughput ceiling and ④ decides the single-request latency floor. ② is the easiest entry point for those with a strong algorithms background (quantization/pruning), ③ is the moat of hardcore systems engineers, and ① is a plus for infrastructure engineers. Pick a route that fits from Learning Paths: Three Routes.
4. Four Perspectives Running Through the Site
- Concept perspective: the core knowledge series explains the principles behind all six layers.
- Case perspective: the eight case studies show how real engines land.
- Paper perspective: paper deep-dives return to the underlying literature.
- Practice perspective: the practice guides walk you through building these layers yourself.
How to use this page
- Want a quick global view? Just read this page through.
- Want to go deep on one layer? Follow the links into the corresponding pages.
- Want to get hands-on? Go straight to Deploy an Inference Service from Scratch and check this six-layer map to see what the demo covers and what it misses — the missing parts are exactly what you should fill in next.
5. One Full Diagram: A Real Production Stack with Six Layers Stacked
Lay out a typical Llama-3-70B production service (vLLM on H100) and the six layers stack like this:
text
Application: OpenAI-compatible API + streaming SSE + Function Calling
Orchestration: vLLM continuous batching + PagedAttention + EAGLE-3 speculative decoding
Engine: vLLM 0.6+ (with RadixAttention prefix sharing)
Operator: FlashAttention-3 + FlashInfer + fused RMSNorm/RoPE kernels
Model: AWQ INT4 weight quantization + FP16 activations (mixed precision)
Hardware: 8x H100 80GB + NVLink + IBOn Llama-3-70B, this stack reaches 5,000-10,000 tokens/s of single-node aggregate throughput — versus HuggingFace Transformers' ~20 tokens/s for a single request, roughly a 250-500x improvement, with not a single weight changed. This is the power of stacking six layers, and the reason inference engineering exists as a discipline.
Further Reading
- What Is Inference Acceleration? — the starting definition of this page
- A Brief History — how the six-layer architecture grew step by step
- Inference vs. Training vs. Fine-Tuning — the uniqueness of inference engineering
- Learning Paths: Three Routes — pick a route by the six layers
- Latency, Throughput, and Concurrency — the inference metric family
- The GPU Memory Hierarchy and the Bandwidth Wall — why LLM inference is memory-bound
- Model Quantization Fundamentals — the core technique of the model layer
- vLLM and PagedAttention — the benchmark of the orchestration layer
- Deploy an Inference Service from Scratch — build all six layers yourself
References
- Kwon et al. Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — the benchmark paper of the orchestration layer
- Dao et al. FlashAttention (NeurIPS 2022) — the representative of the operator layer
- Lin et al. AWQ: Activation-Aware Weight Quantization (MLSys 2024) — the INT4 engineering of the model layer
- NVIDIA. TensorRT-LLM — the engineering reference of the engine layer
- NVIDIA. H100 Tensor Core GPU Architecture White Paper — the foundation of the hardware layer
- Zheng et al. SGLang: Efficient Execution of Structured Language Model Programs (NeurIPS 2024) — prefix sharing and structured acceleration at the orchestration layer