Skip to content

The GPU Memory Hierarchy and the Bandwidth Wall

At a glance LLM inference in the decode phase is throttled almost entirely by GPU memory bandwidth — every generated token requires reading the entire model weights from HBM once. This article covers the memory hierarchy (HBM/L2/SRAM), bandwidth specs across GPU generations, the KV cache memory formula, and how to identify and fight the memory-bound regime.

The GPU Memory Hierarchy and the Bandwidth Wall ​

Concept Definition: Why LLM Inference Is So Memory-Hungry ​

GPU compute (TFLOPS) has grown by orders of magnitude over the years, but memory bandwidth has improved far more slowly. The result: LLM inference in the decode phase can almost never saturate compute — the bottleneck is memory bandwidth. This is the so-called bandwidth wall.

Two key insights for understanding the bandwidth wall:

  1. GPU memory is hierarchical — HBM (large but slow) → L2 cache (small but fast) → shared memory/SRAM (smallest and fastest): the higher you go, the faster and more expensive;
  2. LLM decode is naturally memory-bound — every generated token requires reading the entire model weights from HBM once (weights are too large to fit in SRAM), leaving compute mostly idle.

This is why Weight-Only Quantization and Mixed Precision (compressing weights from FP16 to INT4) can multiply decode throughput several-fold — the bandwidth wall is fought by reducing the number of bytes read. It is also why vLLM and PagedAttention invests so much effort in KV cache memory — the more memory KV cache consumes, the fewer concurrent requests can be accommodated.

1. The Memory Hierarchy ​

GPU storage is a pyramid: higher levels are faster but smaller and more expensive:

text
┌─────────────────────────────┐
│  Registers / SRAM           │  ← Fastest (~20 TB/s), ~256KB/SM
├─────────────────────────────┤
│  L2 Cache                   │  → Fast (~5-12 TB/s), ~40-100MB
├─────────────────────────────┤
│  HBM (High Bandwidth Memory)│  → Slow (~2-5 TB/s), 40-141GB
├─────────────────────────────┤
│  PCIe / Host Memory         │  → Slowest (~30-60 GB/s), effectively unlimited
└─────────────────────────────┘
LevelRoleCapacity (H100 class)BandwidthAccess Latency
Register / SRAMPrivate per-thread, shared per-warp256KB / SM~20 TB/s (on-chip)A few cycles
Shared MemoryShared within a block228KB / SM (configurable)~10 TB/s~10-20 cycles
L2 CacheShared across all SMs50 MB~5-7 TB/s~50-100 cycles
HBMMain GPU memory80 GB~3.35 TB/s~200-400 cycles
Host Memory (CPU DDR)Transferred over PCIeUnlimited~60 GB/s (PCIe4)~1000+ cycles

Why the Hierarchy Matters So Much

CPU/GPU compute is far faster than memory can feed it — the rate at which data moves from HBM to the SM determines whether compute can be saturated. All GPU optimization is, at its core, a fight against data movement: keep hot data in SRAM as much as possible (Kernel Fusion and Custom Kernels), use coalesced access to reduce HBM reads, and use shared-memory tiling to avoid redundant loads. See GPU Architecture and Optimization.

2. Memory and Bandwidth Specs Across GPU Generations ​

When choosing a GPU for LLM inference, the two parameters that matter most are HBM bandwidth (determines decode speed) and HBM capacity (determines how large a model and how much KV cache fit):

GPUMemory TypeCapacityBandwidthBF16 ComputeNotes
A100 80GBHBM2e80 GB~2.0 TB/s312 TFLast-gen flagship, still mainstream
A100 40GBHBM2e40 GB~1.55 TB/s312 TFSmaller-memory variant of the same GPU
H100 SXM5HBM380 GB~3.35 TB/s989 TFCurrent workhorse
H100 PCIeHBM380 GB~2 TB/s989 TFPCIe variant with reduced bandwidth
H200 SXMHBM3e141 GB~4.8 TB/s989 TFLarge memory + high bandwidth
H20HBM396 GB~4 TB/s148 TF (FP16)Compliance variant with capped compute
L40SGDDR648 GB~0.866 TB/s362 TFInference-optimized, good value
B200 SXMHBM3e192 GB~8 TB/s2.25 PF (FP4)Blackwell flagship
B100HBM3e192 GB~8 TB/s1.8 PFMid-range Blackwell

A Simple Decision Guide for Choosing a GPU

  • Generous budget + extreme latency → H200 (4.8 TB/s bandwidth + 141GB memory — the LLM inference sweet spot)
  • High-concurrency online serving → H100 SXM5 (enough bandwidth, mature ecosystem)
  • Offline batch processing → A100 80GB (still good value)
  • Not enough memory? Quantize → W4A16 quantization squeezes a 70B model onto a single H100 80GB

More selection details in the Hardware Primer and Benchmark Data & Tool Profiles.

3. The Bandwidth Wall: The Core Bottleneck of Decode ​

Do the Math ​

Take Llama-2-70B (70B parameters, FP16 weights ≈ 140GB):

  • In decode, every generated token requires reading all weights from HBM to the SM at least once;
  • Total weights 140GB at H100 bandwidth of 3.35TB/s → theoretical best ~42ms/token;
  • That's ≈ 24 tokens/s (single request, single GPU ceiling).

Meanwhile, H100's BF16 compute is 989 TFLOPS, and a 70B model needs roughly 140 GFLOPS per token — compute can finish a token in just 0.14ms. Compute and bandwidth differ by 300×. That is the essence of the bandwidth wall:

text
Compute-limited latency:    0.14 ms/token  (989 TFLOPS ÷ 140 GFLOPS)
Bandwidth-limited latency:  42   ms/token  (140GB ÷ 3.35 TB/s)
Actual latency:             45-50 ms/token (bandwidth wall dominates, 99% of compute idle)

Don't Be Fooled by "99% of Compute Idle"

This is a structural fact of LLM inference, not a sign of poor optimization — single-request decode simply cannot saturate compute. The only remedy is batching: 32 concurrent requests decoding together amortize compute 32-fold, while each token still reads the weights only once (weights are shared across requests), instantly multiplying throughput 30×. This is the principle behind Batching and Request Scheduling's continuous batching.

The Arithmetic Intensity of Decode ​

A more precise analysis uses The Roofline Model and Compute Analysis — an operator's arithmetic intensity (AI) = FLOPS / Bytes. Per token in decode:

  • FLOPS: each weight participates in 2 operations (multiply-add) ≈ 140 GFLOPS;
  • Bytes: reading 140GB of weights from HBM ≈ 280 GB (FP16, 2 bytes per parameter);
  • AI = 140 / 280 ≈ 0.5 FLOPS/Byte.

The H100 ridge point ≈ compute / bandwidth = 989 TFLOPS / 3.35 TB/s ≈ 295 FLOPS/Byte. Decode's AI is far below the ridge → strongly memory-bound. By contrast, large-batch matmul can reach an AI in the hundreds — compute-bound.

4. KV Cache: The Other Big Consumer of Memory ​

Beyond weights, LLM inference must reserve a KV cache for every active request — every generated token queries the K and V vectors of all historical tokens. The KV cache memory formula:

text
KV_cache_bytes = 2  ×  seq_len  ×  batch  ×  num_layers  ×  num_heads  ×  head_dim  ×  dtype_bytes
                 ↑      ↑          ↑          ↑                ↑             ↑           ↑
              K & V   length    requests   #layers         #heads        head dim    FP16=2,INT8=1

Take Llama-2-70B (80 layers, 64 heads, 128 dims per head) as an example, in FP16 per request:

text
2 × 2048 × 1 × 80 × 64 × 128 × 2 = 5.4 GB / request

A single 2048-token request occupies 5.4GB of memory! Ten concurrent requests take 54GB — compared to 140GB of weights, KV cache eats a substantial chunk of memory, directly limiting concurrency.

Several Ways to Compress the KV Cache

  1. GQA / MQA: grouped-query attention reduces num_heads by 4-8× (Llama-2-70B uses 8 GQA groups, shrinking the KV cache to 1/8 directly);
  2. KV cache quantization: compress KV from FP16 to INT8 or FP8, nearly lossless (Model Quantization Fundamentals);
  3. PagedAttention: vLLM and PagedAttention manages KV cache with OS-style paging, eliminating fragmentation and lifting memory utilization from ~60% to ~95%;
  4. Prefix caching: requests sharing the same prompt prefix reuse the KV cache — huge memory savings for long system prompts.

5. Memory-Bound vs. Compute-Bound: How to Tell ​

To determine whether an operator or phase is memory-bound or compute-bound, use the arithmetic intensity AI from The Roofline Model and Compute Analysis:

OperatorAI (FLOPS/Byte)TypeNotes
Elementwise (add, relu, scale)~1memory-bound1 byte read per 1 operation
LayerNorm~2-5memory-boundReduce + multiple read/write passes
Activation (GELU, Softmax)~1-5memory-boundSame as above
Vector-vector op (small vector add)~1memory-boundReads far more than it computes
matmul (small-batch decode)~1-10memory-boundTypical of the decode phase
matmul (large-batch prefill)Hundredscompute-boundThe prefill phase
Tensor Core matmul (M=N=K=8192)Several thousandcompute-boundNearly saturates compute

A Simple Rule of Thumb

Operators that do too little math (elementwise, small-batch matmul) → memory-bound; operators that do math at full tilt (large GEMM) → compute-bound. Optimization directions:

6. Engineering Weapons Against the Bandwidth Wall ​

The bandwidth wall is not unbreakable — you just need the right tools:

TechniquePrincipleGainDetails
Weight quantization (W4A16)4-bit weights → 4× fewer bytes to read~3-4× decode throughputWeight-Only Quantization and Mixed Precision
KV cache quantizationKV in FP8/INT8Doubles memory → doubles concurrencyModel Quantization Fundamentals
Kernel fusionFuse multiple elementwise ops into one kernel, read fewer timesFewer HBM round tripsKernel Fusion and Custom Kernels
FlashAttentionKeep attention intermediates in SRAM, never write back to HBMMemory/speed win-win for long contextKernel Fusion and Custom Kernels
PagedAttentionPaged KV management, no fragmentation+30-60% usable concurrencyvLLM and PagedAttention
Continuous batchingPack decode batches full to drain bandwidth~30× single-GPU throughputBatching and Request Scheduling
Speculative decodingA small model proposes candidates; the large model verifies in batchFewer decode stepsSpeculative Decoding and Medusa/EAGLE

Suggested order of attack: start with Model Quantization Fundamentals (cheapest) → then Kernel Fusion and Custom Kernels / FlashAttention (built into inference engines) → finally Batching and Request Scheduling (deploy vLLM).

7. Trade-offs ​

  • Bandwidth vs. compute: choose GPUs based on your bottleneck. LLM inference favors high-bandwidth cards (H100/H200); training favors high-compute cards (B200); edge deployment favors value (L40S);
  • Memory capacity vs. bandwidth: H200 141GB > H100 80GB, but bandwidth is only 1.4× higher — large models prefer H200, while small models run more efficiently on H100;
  • Precision vs. speed: FP16 → INT8 → INT4 — each step halves bandwidth pressure, but accuracy suffers — W4A16 from Weight-Only Quantization and Mixed Precision is the current sweet spot;
  • Hardware vs. software: buy H200, or spend time on vLLM + AWQ? Small teams should optimize software first, then consider a hardware upgrade.

Further Reading ​

References ​