Appearance
The Roofline Model and Compute Analysis
Concept Definition: Seeing Through Operator Performance in One Chart
The Roofline model (Williams et al., 2009) is a unified framework for analyzing operator performance on a processor — in one sentence:
text
Attainable FLOPS = min( Peak compute , bandwidth × arithmetic intensity )Every operator can be plotted as a single point on the Roofline chart. Its position tells you: whether this operator is held back by compute or by bandwidth, which in turn dictates the optimization direction. Ninety percent of LLM inference optimization decisions start with drawing a Roofline chart — it turns "why is this slow" from guesswork into a quantitative question.
Three key insights for understanding Roofline:
- Arithmetic intensity (AI) = FLOPS / Bytes — how many operations you get out of each byte read. High AI → compute-intensive; low AI → memory-intensive;
- Every piece of hardware has a "roofline" — a sloped segment on the left (bandwidth limit) and a horizontal segment on the right (compute limit), with a ridge point in between;
- An operator's position determines its bottleneck — AI left of the ridge point → memory-bound; right → compute-bound.
1. Arithmetic Intensity: An Operator's "Value for Money"
Arithmetic intensity (also called the ops-to-bytes ratio) is the core quantity of Roofline:
text
AI = total FLOPS / total bytes moved- Numerator: how many floating-point operations the operator performs (2MNK for matmul, N for elementwise);
- Denominator: how many bytes the operator moves from/to HBM (input + output).
AI of Common Operators in LLM Inference
| Operator | FLOPs | Bytes Moved | Arithmetic Intensity | Type |
|---|---|---|---|---|
| Elementwise add/scale | N | 3N (2 inputs + 1 output) | ~0.33 | memory-bound |
| ReLU/GELU | ~N | 2N (input + output) | ~0.5 | memory-bound |
| LayerNorm | ~5N | 2N | ~2.5 | memory-bound |
| Softmax (row-wise) | ~5N | 2N | ~2.5 | memory-bound |
| Small-batch decode matmul | 2·M·N·K (M=1) | 2·N·K + 2·M·K | ~1-5 | memory-bound |
| Large-batch prefill matmul | 2·M·N·K (large M) | 2·N·K + 2·M·K | Tens to hundreds | compute-bound |
| Large GEMM (M=N=K=8192, FP16) | ~10⁹ | ~10⁷ | Hundreds to thousands | compute-bound |
Why Decode Matmul Is Memory-Bound
In decode, only one token is generated at a time: the M dimension (batch × 1) = 1, so the matmul degenerates into a "vector × matrix" product. Compute does very little (M=1), yet the entire N×K weight matrix must be read from HBM — compute sits idle while bandwidth is saturated. See The GPU Memory Hierarchy and the Bandwidth Wall.
2. The Roofline Chart: The Hardware "Performance Ceiling"
Every GPU has its own roofline. Take the H100 SXM5 as an example:
- Peak compute (BF16, Tensor Core) ≈ 989 TFLOPS = 989 × 10¹² FLOPS/s
- HBM bandwidth ≈ 3.35 TB/s = 3.35 × 10¹² Bytes/s
- Ridge-point AI = 989 / 3.35 ≈ 295 FLOPS/Byte
Drawn as a chart:
text
Attainable compute (TFLOPS)
↑
1000┤ ╱───────────────── ← compute ceiling (989 TFLOPS)
900┤ ╱──
800┤ ╱──
700┤ ╱──
600┤ ╱──
500┤ ╱──
400┤ ╱──
300┤ ╱── ← ridge point AI ≈ 295 FLOPS/Byte
200┤ ╱──
100┤╱─
0┼────────────────────────────────────────→ Arithmetic intensity (FLOPS/Byte)
0 50 100 150 200 300 500 1000
(log scale)Reading the chart:
- Left half (AI < 295): attainable compute = bandwidth × AI — the sloped segment (bandwidth-limited). The lower the AI, the less compute is utilized;
- Right half (AI > 295): attainable compute = peak compute — the horizontal segment (compute-limited). The operator's compute is pinned against the ceiling;
- Ridge point (AI = 295): where the slope meets the horizontal line — the critical point where the bandwidth limit gives way to the compute limit.
How to Read a Roofline Chart
- Compute the operator's AI: FLOPS / Bytes;
- Mark that AI on the horizontal axis;
- Read the attainable compute straight up from that point — this is the operator's theoretical ceiling on this GPU;
- Measured vs. theoretical: if the measured value is far below the ceiling → that's where the optimization headroom is; if the measured value is close to the ceiling → you've hit the top and must change hardware or algorithm.
3. Roofline Parameters Across GPU Generations
Different GPUs have ridge points in different positions — compute grows fast while bandwidth grows slowly, so each new generation shifts the ridge point to the right, turning more operators memory-bound:
| GPU | Compute (BF16) | Bandwidth | Ridge-point AI |
|---|---|---|---|
| A100 80GB | 312 TF | 2.0 TB/s | 156 |
| H100 SXM5 | 989 TF | 3.35 TB/s | 295 |
| H200 SXM | 989 TF | 4.8 TB/s | 206 |
| B200 SXM | 2250 TF (FP4) | 8 TB/s | 281 |
The Consequence of "Fast Compute, Slow Bandwidth"
A100 → H100: compute grew 3×, bandwidth only 1.7×. The result: decode matmul on the H100 is even more memory-bound than on the A100 — upgrading the GPU alone won't make decode 3× faster; it may only be 1.7× faster (bandwidth caps it). This is why the H200 actually suits LLM inference better than the H100 — same compute (989 TF), but bandwidth jumps from 3.35 to 4.8 TB/s, directly relieving the decode bandwidth wall.
4. Diagnosing LLM Inference with Roofline
1. Decode Phase
Llama-2-70B on a single H100 SXM5, single-request decode:
- Compute per token: ~140 GFLOPS (70B parameters × 2 FLOPS/parameter);
- Bandwidth per token: 140GB weights × 2 bytes = 280GB (FP16);
- AI = 140 / 280 ≈ 0.5 FLOPS/Byte;
- On the Roofline, AI = 0.5 → attainable compute = 3.35TB/s × 0.5 = 1.675 TFLOPS;
- Actual compute (at 50ms/token) ≈ 140GFLOPS / 50ms = 2.8 TFLOPS;
- Attainable ceiling 1.675 TFLOPS; compute utilization 0.3% (the 989-TFLOPS ceiling is nowhere near).
Diagnosis: strongly memory-bound — more GPU compute won't help; bandwidth pressure must be reduced — apply Weight-Only Quantization and Mixed Precision (W4A16, turning 280GB → 70GB; AI stays the same but bandwidth pressure drops 4×) and increase batch size (multiple requests share the weight reads, so AI rises linearly with batch).
2. Prefill Phase
Llama-2-70B on a single H100 SXM5, single-request prefill with a 2048-token prompt:
- Compute: ~2 × 2048 × 70B FLOPS ≈ 286 TFLOPS;
- Bandwidth: ~140GB (weights read once; prompt input is negligible);
- AI = 286 × 10¹² / 280 × 10⁹ ≈ 1020 FLOPS/Byte;
- On the Roofline, AI = 1020 exceeds the ridge point of 295 → compute-bound;
- Attainable compute = 989 TFLOPS (pinned against the compute ceiling).
Diagnosis: prefill is compute-bound, so adding compute (upgrading to B200) genuinely works. This is why vLLM and PagedAttention schedules prefill and decode separately — they have different bottlenecks and must be optimized separately.
3. Elementwise Operators
LayerNorm has AI ≈ 2.5, far below the ridge point of 295:
- Attainable compute = 3.35TB/s × 2.5 ≈ 8.4 TFLOPS (0.8% of 989 TFLOPS);
- If not fused, measured performance reads the input and output multiple times — fusing 3-4 elementwise operators into one kernel with Kernel Fusion and Custom Kernels leaves AI unchanged but cuts read/write passes from 4× to 1× — effective bandwidth ×4.
5. Choosing Optimization Strategies with Roofline
Roofline gives you not just a diagnosis but a priority order for optimization:
| Operator Position | AI | Bottleneck | Optimization Direction |
|---|---|---|---|
| Far left (AI ≈ 0.5) | Extremely low | Bandwidth | Quantization to cut bytes, larger batch, kernel fusion |
| Mid-left (AI ~50) | Low-medium | Bandwidth | Kernel fusion, loop-order tuning, tiling |
| Near the ridge (AI ~295) | Critical | Both | Optimizable from either side — check how far measured performance sits from the Roofline |
| Just right of the ridge (AI ~500) | Medium-high | Compute | Reduce computation (distillation, pruning), use Tensor Cores |
| Far right (AI ~2000) | High | Compute | Higher-compute hardware, lower precision (FP8/FP4) |
A General Optimization Mantra
- Memory-bound operators: first cut bytes (Model Quantization Fundamentals), then cut passes (Kernel Fusion and Custom Kernels), and finally enlarge batch (Batching and Request Scheduling);
- Compute-bound operators: first cut FLOPs (Knowledge Distillation, Pruning and Sparsification), then switch to higher-throughput precision (FP8/FP4), and finally upgrade hardware.
6. Limitations of Roofline
Roofline is a simplified model; real operators are also affected by:
- Cache hits: Roofline assumes all data comes from HBM, but L2 hits make effective bandwidth much higher (even approaching SRAM bandwidth);
- Intrinsic operator parallelism: a matmul running on Tensor Cores vs. plain CUDA Cores differs by tens of times — Roofline doesn't distinguish;
- Instruction mix: an operator mixing add/sub/mul/div may not reach peak utilization (Tensor Cores only favor dense GEMM);
- Memory access pattern: coalesced vs. random access differs by 5-10× in bandwidth — Roofline counts totals but doesn't look at patterns;
- Latency hiding: GPUs hide HBM latency through multi-warp concurrency, but bandwidth still can't be saturated when there aren't enough warps.
Roofline Gives an Upper Bound, Not a Measurement
The Roofline gives a theoretical ceiling — if your operator measures at only 30% of its Roofline ceiling, the missing 70% comes from other factors (cache, warp occupancy, bank conflicts, instruction mix). To chase that, use Nsight Compute for single-operator profiling. See GPU Architecture and Optimization and Inference Benchmarking in Practice.
7. Trade-offs
- Theory vs. measurement: Roofline is the starting point — it tells you "whether to optimize compute or bandwidth," but the actual gain must be measured;
- Compute vs. bandwidth when choosing hardware: compute-bound tasks (training) favor high-compute cards (B200); memory-bound tasks (LLM inference) favor high-bandwidth cards (H200);
- Quantization vs. distillation: decide by operator position — memory-bound favors quantization first (directly cuts bytes); compute-bound calls for distillation to cut computation;
- Budget allocation: small teams should first reach the Roofline ceiling in software (vLLM + AWQ + kernel fusion) before considering hardware upgrades — in many cases software alone captures 80% of the gain.
Further Reading
- The GPU Memory Hierarchy and the Bandwidth Wall — the causes of the bandwidth wall and the hardware data
- Weight-Only Quantization and Mixed Precision — cutting bytes directly to break through the bandwidth wall
- Kernel Fusion and Custom Kernels — reducing the number of HBM read/write passes
- GPU Architecture and Optimization — the low-level view of cache/occupancy/bank conflicts
- Batching and Request Scheduling — enlarging batch to raise matmul AI
- vLLM and PagedAttention — the industrial implementation of Roofline-driven optimization
- Inference Benchmarking in Practice — putting the Roofline model into practice
References
- Williams, Waterman, Patterson. Roofline: An Insightful Visual Performance Model for Multicore Architectures (CACM 2009) — the original Roofline paper
- Ofenbeck et al. Roofline Model for GPUs (2017) — extending Roofline to GPUs
- NVIDIA Nsight Compute Documentation — single-operator profiling tool
- Ivanov et al. Data Movement Is All You Need (NeurIPS 2021) — analyzing bandwidth bottlenecks of neural network operators with Roofline
- Pope et al. Efficiently Scaling Transformer Inference (MLSys 2023) — the systematic application of Roofline to LLM inference