Appearance
The GPU Memory Hierarchy and the Bandwidth Wall
Concept Definition: Why LLM Inference Is So Memory-Hungry
GPU compute (TFLOPS) has grown by orders of magnitude over the years, but memory bandwidth has improved far more slowly. The result: LLM inference in the decode phase can almost never saturate compute — the bottleneck is memory bandwidth. This is the so-called bandwidth wall.
Two key insights for understanding the bandwidth wall:
- GPU memory is hierarchical — HBM (large but slow) → L2 cache (small but fast) → shared memory/SRAM (smallest and fastest): the higher you go, the faster and more expensive;
- LLM decode is naturally memory-bound — every generated token requires reading the entire model weights from HBM once (weights are too large to fit in SRAM), leaving compute mostly idle.
This is why Weight-Only Quantization and Mixed Precision (compressing weights from FP16 to INT4) can multiply decode throughput several-fold — the bandwidth wall is fought by reducing the number of bytes read. It is also why vLLM and PagedAttention invests so much effort in KV cache memory — the more memory KV cache consumes, the fewer concurrent requests can be accommodated.
1. The Memory Hierarchy
GPU storage is a pyramid: higher levels are faster but smaller and more expensive:
text
┌─────────────────────────────┐
│ Registers / SRAM │ ← Fastest (~20 TB/s), ~256KB/SM
├─────────────────────────────┤
│ L2 Cache │ → Fast (~5-12 TB/s), ~40-100MB
├─────────────────────────────┤
│ HBM (High Bandwidth Memory)│ → Slow (~2-5 TB/s), 40-141GB
├─────────────────────────────┤
│ PCIe / Host Memory │ → Slowest (~30-60 GB/s), effectively unlimited
└─────────────────────────────┘| Level | Role | Capacity (H100 class) | Bandwidth | Access Latency |
|---|---|---|---|---|
| Register / SRAM | Private per-thread, shared per-warp | 256KB / SM | ~20 TB/s (on-chip) | A few cycles |
| Shared Memory | Shared within a block | 228KB / SM (configurable) | ~10 TB/s | ~10-20 cycles |
| L2 Cache | Shared across all SMs | 50 MB | ~5-7 TB/s | ~50-100 cycles |
| HBM | Main GPU memory | 80 GB | ~3.35 TB/s | ~200-400 cycles |
| Host Memory (CPU DDR) | Transferred over PCIe | Unlimited | ~60 GB/s (PCIe4) | ~1000+ cycles |
Why the Hierarchy Matters So Much
CPU/GPU compute is far faster than memory can feed it — the rate at which data moves from HBM to the SM determines whether compute can be saturated. All GPU optimization is, at its core, a fight against data movement: keep hot data in SRAM as much as possible (Kernel Fusion and Custom Kernels), use coalesced access to reduce HBM reads, and use shared-memory tiling to avoid redundant loads. See GPU Architecture and Optimization.
2. Memory and Bandwidth Specs Across GPU Generations
When choosing a GPU for LLM inference, the two parameters that matter most are HBM bandwidth (determines decode speed) and HBM capacity (determines how large a model and how much KV cache fit):
| GPU | Memory Type | Capacity | Bandwidth | BF16 Compute | Notes |
|---|---|---|---|---|---|
| A100 80GB | HBM2e | 80 GB | ~2.0 TB/s | 312 TF | Last-gen flagship, still mainstream |
| A100 40GB | HBM2e | 40 GB | ~1.55 TB/s | 312 TF | Smaller-memory variant of the same GPU |
| H100 SXM5 | HBM3 | 80 GB | ~3.35 TB/s | 989 TF | Current workhorse |
| H100 PCIe | HBM3 | 80 GB | ~2 TB/s | 989 TF | PCIe variant with reduced bandwidth |
| H200 SXM | HBM3e | 141 GB | ~4.8 TB/s | 989 TF | Large memory + high bandwidth |
| H20 | HBM3 | 96 GB | ~4 TB/s | 148 TF (FP16) | Compliance variant with capped compute |
| L40S | GDDR6 | 48 GB | ~0.866 TB/s | 362 TF | Inference-optimized, good value |
| B200 SXM | HBM3e | 192 GB | ~8 TB/s | 2.25 PF (FP4) | Blackwell flagship |
| B100 | HBM3e | 192 GB | ~8 TB/s | 1.8 PF | Mid-range Blackwell |
A Simple Decision Guide for Choosing a GPU
- Generous budget + extreme latency → H200 (4.8 TB/s bandwidth + 141GB memory — the LLM inference sweet spot)
- High-concurrency online serving → H100 SXM5 (enough bandwidth, mature ecosystem)
- Offline batch processing → A100 80GB (still good value)
- Not enough memory? Quantize → W4A16 quantization squeezes a 70B model onto a single H100 80GB
More selection details in the Hardware Primer and Benchmark Data & Tool Profiles.
3. The Bandwidth Wall: The Core Bottleneck of Decode
Do the Math
Take Llama-2-70B (70B parameters, FP16 weights ≈ 140GB):
- In decode, every generated token requires reading all weights from HBM to the SM at least once;
- Total weights 140GB at H100 bandwidth of 3.35TB/s → theoretical best ~42ms/token;
- That's ≈ 24 tokens/s (single request, single GPU ceiling).
Meanwhile, H100's BF16 compute is 989 TFLOPS, and a 70B model needs roughly 140 GFLOPS per token — compute can finish a token in just 0.14ms. Compute and bandwidth differ by 300×. That is the essence of the bandwidth wall:
text
Compute-limited latency: 0.14 ms/token (989 TFLOPS ÷ 140 GFLOPS)
Bandwidth-limited latency: 42 ms/token (140GB ÷ 3.35 TB/s)
Actual latency: 45-50 ms/token (bandwidth wall dominates, 99% of compute idle)Don't Be Fooled by "99% of Compute Idle"
This is a structural fact of LLM inference, not a sign of poor optimization — single-request decode simply cannot saturate compute. The only remedy is batching: 32 concurrent requests decoding together amortize compute 32-fold, while each token still reads the weights only once (weights are shared across requests), instantly multiplying throughput 30×. This is the principle behind Batching and Request Scheduling's continuous batching.
The Arithmetic Intensity of Decode
A more precise analysis uses The Roofline Model and Compute Analysis — an operator's arithmetic intensity (AI) = FLOPS / Bytes. Per token in decode:
- FLOPS: each weight participates in 2 operations (multiply-add) ≈ 140 GFLOPS;
- Bytes: reading 140GB of weights from HBM ≈ 280 GB (FP16, 2 bytes per parameter);
- AI = 140 / 280 ≈ 0.5 FLOPS/Byte.
The H100 ridge point ≈ compute / bandwidth = 989 TFLOPS / 3.35 TB/s ≈ 295 FLOPS/Byte. Decode's AI is far below the ridge → strongly memory-bound. By contrast, large-batch matmul can reach an AI in the hundreds — compute-bound.
4. KV Cache: The Other Big Consumer of Memory
Beyond weights, LLM inference must reserve a KV cache for every active request — every generated token queries the K and V vectors of all historical tokens. The KV cache memory formula:
text
KV_cache_bytes = 2 × seq_len × batch × num_layers × num_heads × head_dim × dtype_bytes
↑ ↑ ↑ ↑ ↑ ↑ ↑
K & V length requests #layers #heads head dim FP16=2,INT8=1Take Llama-2-70B (80 layers, 64 heads, 128 dims per head) as an example, in FP16 per request:
text
2 × 2048 × 1 × 80 × 64 × 128 × 2 = 5.4 GB / requestA single 2048-token request occupies 5.4GB of memory! Ten concurrent requests take 54GB — compared to 140GB of weights, KV cache eats a substantial chunk of memory, directly limiting concurrency.
Several Ways to Compress the KV Cache
- GQA / MQA: grouped-query attention reduces num_heads by 4-8× (Llama-2-70B uses 8 GQA groups, shrinking the KV cache to 1/8 directly);
- KV cache quantization: compress KV from FP16 to INT8 or FP8, nearly lossless (Model Quantization Fundamentals);
- PagedAttention: vLLM and PagedAttention manages KV cache with OS-style paging, eliminating fragmentation and lifting memory utilization from ~60% to ~95%;
- Prefix caching: requests sharing the same prompt prefix reuse the KV cache — huge memory savings for long system prompts.
5. Memory-Bound vs. Compute-Bound: How to Tell
To determine whether an operator or phase is memory-bound or compute-bound, use the arithmetic intensity AI from The Roofline Model and Compute Analysis:
| Operator | AI (FLOPS/Byte) | Type | Notes |
|---|---|---|---|
| Elementwise (add, relu, scale) | ~1 | memory-bound | 1 byte read per 1 operation |
| LayerNorm | ~2-5 | memory-bound | Reduce + multiple read/write passes |
| Activation (GELU, Softmax) | ~1-5 | memory-bound | Same as above |
| Vector-vector op (small vector add) | ~1 | memory-bound | Reads far more than it computes |
| matmul (small-batch decode) | ~1-10 | memory-bound | Typical of the decode phase |
| matmul (large-batch prefill) | Hundreds | compute-bound | The prefill phase |
| Tensor Core matmul (M=N=K=8192) | Several thousand | compute-bound | Nearly saturates compute |
A Simple Rule of Thumb
Operators that do too little math (elementwise, small-batch matmul) → memory-bound; operators that do math at full tilt (large GEMM) → compute-bound. Optimization directions:
- memory-bound → reduce bytes read (Kernel Fusion and Custom Kernels, Model Quantization Fundamentals to compress weights);
- compute-bound → reduce computation (Knowledge Distillation, Pruning and Sparsification to cut parameters).
6. Engineering Weapons Against the Bandwidth Wall
The bandwidth wall is not unbreakable — you just need the right tools:
| Technique | Principle | Gain | Details |
|---|---|---|---|
| Weight quantization (W4A16) | 4-bit weights → 4× fewer bytes to read | ~3-4× decode throughput | Weight-Only Quantization and Mixed Precision |
| KV cache quantization | KV in FP8/INT8 | Doubles memory → doubles concurrency | Model Quantization Fundamentals |
| Kernel fusion | Fuse multiple elementwise ops into one kernel, read fewer times | Fewer HBM round trips | Kernel Fusion and Custom Kernels |
| FlashAttention | Keep attention intermediates in SRAM, never write back to HBM | Memory/speed win-win for long context | Kernel Fusion and Custom Kernels |
| PagedAttention | Paged KV management, no fragmentation | +30-60% usable concurrency | vLLM and PagedAttention |
| Continuous batching | Pack decode batches full to drain bandwidth | ~30× single-GPU throughput | Batching and Request Scheduling |
| Speculative decoding | A small model proposes candidates; the large model verifies in batch | Fewer decode steps | Speculative Decoding and Medusa/EAGLE |
Suggested order of attack: start with Model Quantization Fundamentals (cheapest) → then Kernel Fusion and Custom Kernels / FlashAttention (built into inference engines) → finally Batching and Request Scheduling (deploy vLLM).
7. Trade-offs
- Bandwidth vs. compute: choose GPUs based on your bottleneck. LLM inference favors high-bandwidth cards (H100/H200); training favors high-compute cards (B200); edge deployment favors value (L40S);
- Memory capacity vs. bandwidth: H200 141GB > H100 80GB, but bandwidth is only 1.4× higher — large models prefer H200, while small models run more efficiently on H100;
- Precision vs. speed: FP16 → INT8 → INT4 — each step halves bandwidth pressure, but accuracy suffers — W4A16 from Weight-Only Quantization and Mixed Precision is the current sweet spot;
- Hardware vs. software: buy H200, or spend time on vLLM + AWQ? Small teams should optimize software first, then consider a hardware upgrade.
Further Reading
- The Roofline Model and Compute Analysis — quantitatively classifying memory-bound vs. compute-bound with AI
- Weight-Only Quantization and Mixed Precision — the first weapon against the bandwidth wall
- Kernel Fusion and Custom Kernels — how FlashAttention avoids HBM round trips
- Batching and Request Scheduling — piling on concurrency to drain bandwidth
- GPU Architecture and Optimization — the low-level view of SM/warp/shared memory
- vLLM and PagedAttention — how PagedAttention solves KV cache fragmentation
- Hardware Primer — GPU selection and configuration
References
- NVIDIA H100 White Paper — official H100 architecture and HBM3 specs
- NVIDIA H200 Datasheet — H200 HBM3e 141GB specs
- NVIDIA H100 Memory Bandwidth Analysis (Markomanlicious blog) — measured data for each level of the memory hierarchy
- Pope et al. Efficiently Scaling Transformer Inference (MLSys 2023) — a systematic analysis of the LLM inference memory wall
- Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023) — KV cache memory management and PagedAttention
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention (NeurIPS 2022) — the landmark algorithm that broke through the HBM bandwidth wall