Skip to content

The Roofline Model and Compute Analysis

At a glance The Roofline model is a unified framework for analyzing GPU operator performance — attainable FLOPS = min(peak FLOPS, bandwidth × arithmetic_intensity). This article draws a Roofline chart, explains arithmetic intensity, the ridge point, and the memory-bound vs. compute-bound classification, and uses it to diagnose LLM inference bottlenecks.

The Roofline Model and Compute Analysis ​

Concept Definition: Seeing Through Operator Performance in One Chart ​

The Roofline model (Williams et al., 2009) is a unified framework for analyzing operator performance on a processor — in one sentence:

text
Attainable FLOPS = min( Peak compute , bandwidth × arithmetic intensity )

Every operator can be plotted as a single point on the Roofline chart. Its position tells you: whether this operator is held back by compute or by bandwidth, which in turn dictates the optimization direction. Ninety percent of LLM inference optimization decisions start with drawing a Roofline chart — it turns "why is this slow" from guesswork into a quantitative question.

Three key insights for understanding Roofline:

  1. Arithmetic intensity (AI) = FLOPS / Bytes — how many operations you get out of each byte read. High AI → compute-intensive; low AI → memory-intensive;
  2. Every piece of hardware has a "roofline" — a sloped segment on the left (bandwidth limit) and a horizontal segment on the right (compute limit), with a ridge point in between;
  3. An operator's position determines its bottleneck — AI left of the ridge point → memory-bound; right → compute-bound.

1. Arithmetic Intensity: An Operator's "Value for Money" ​

Arithmetic intensity (also called the ops-to-bytes ratio) is the core quantity of Roofline:

text
AI = total FLOPS / total bytes moved
  • Numerator: how many floating-point operations the operator performs (2MNK for matmul, N for elementwise);
  • Denominator: how many bytes the operator moves from/to HBM (input + output).

AI of Common Operators in LLM Inference ​

OperatorFLOPsBytes MovedArithmetic IntensityType
Elementwise add/scaleN3N (2 inputs + 1 output)~0.33memory-bound
ReLU/GELU~N2N (input + output)~0.5memory-bound
LayerNorm~5N2N~2.5memory-bound
Softmax (row-wise)~5N2N~2.5memory-bound
Small-batch decode matmul2·M·N·K (M=1)2·N·K + 2·M·K~1-5memory-bound
Large-batch prefill matmul2·M·N·K (large M)2·N·K + 2·M·KTens to hundredscompute-bound
Large GEMM (M=N=K=8192, FP16)~10⁹~10⁷Hundreds to thousandscompute-bound

Why Decode Matmul Is Memory-Bound

In decode, only one token is generated at a time: the M dimension (batch × 1) = 1, so the matmul degenerates into a "vector × matrix" product. Compute does very little (M=1), yet the entire N×K weight matrix must be read from HBM — compute sits idle while bandwidth is saturated. See The GPU Memory Hierarchy and the Bandwidth Wall.

2. The Roofline Chart: The Hardware "Performance Ceiling" ​

Every GPU has its own roofline. Take the H100 SXM5 as an example:

  • Peak compute (BF16, Tensor Core) ≈ 989 TFLOPS = 989 × 10¹² FLOPS/s
  • HBM bandwidth ≈ 3.35 TB/s = 3.35 × 10¹² Bytes/s
  • Ridge-point AI = 989 / 3.35 ≈ 295 FLOPS/Byte

Drawn as a chart:

text
Attainable compute (TFLOPS)
  ↑
1000┤                                  ╱─────────────────  ← compute ceiling (989 TFLOPS)
 900┤                              ╱──
 800┤                          ╱──
 700┤                      ╱──
 600┤                  ╱──
 500┤              ╱──
 400┤          ╱──
 300┤      ╱──  ← ridge point AI ≈ 295 FLOPS/Byte
 200┤  ╱──
 100┤╱─
   0┼────────────────────────────────────────→ Arithmetic intensity (FLOPS/Byte)
    0    50   100  150  200  300  500  1000
                  (log scale)

Reading the chart:

  • Left half (AI < 295): attainable compute = bandwidth × AI — the sloped segment (bandwidth-limited). The lower the AI, the less compute is utilized;
  • Right half (AI > 295): attainable compute = peak compute — the horizontal segment (compute-limited). The operator's compute is pinned against the ceiling;
  • Ridge point (AI = 295): where the slope meets the horizontal line — the critical point where the bandwidth limit gives way to the compute limit.

How to Read a Roofline Chart

  1. Compute the operator's AI: FLOPS / Bytes;
  2. Mark that AI on the horizontal axis;
  3. Read the attainable compute straight up from that point — this is the operator's theoretical ceiling on this GPU;
  4. Measured vs. theoretical: if the measured value is far below the ceiling → that's where the optimization headroom is; if the measured value is close to the ceiling → you've hit the top and must change hardware or algorithm.

3. Roofline Parameters Across GPU Generations ​

Different GPUs have ridge points in different positions — compute grows fast while bandwidth grows slowly, so each new generation shifts the ridge point to the right, turning more operators memory-bound:

GPUCompute (BF16)BandwidthRidge-point AI
A100 80GB312 TF2.0 TB/s156
H100 SXM5989 TF3.35 TB/s295
H200 SXM989 TF4.8 TB/s206
B200 SXM2250 TF (FP4)8 TB/s281

The Consequence of "Fast Compute, Slow Bandwidth"

A100 → H100: compute grew 3×, bandwidth only 1.7×. The result: decode matmul on the H100 is even more memory-bound than on the A100 — upgrading the GPU alone won't make decode 3× faster; it may only be 1.7× faster (bandwidth caps it). This is why the H200 actually suits LLM inference better than the H100 — same compute (989 TF), but bandwidth jumps from 3.35 to 4.8 TB/s, directly relieving the decode bandwidth wall.

4. Diagnosing LLM Inference with Roofline ​

1. Decode Phase ​

Llama-2-70B on a single H100 SXM5, single-request decode:

  • Compute per token: ~140 GFLOPS (70B parameters × 2 FLOPS/parameter);
  • Bandwidth per token: 140GB weights × 2 bytes = 280GB (FP16);
  • AI = 140 / 280 ≈ 0.5 FLOPS/Byte;
  • On the Roofline, AI = 0.5 → attainable compute = 3.35TB/s × 0.5 = 1.675 TFLOPS;
  • Actual compute (at 50ms/token) ≈ 140GFLOPS / 50ms = 2.8 TFLOPS;
  • Attainable ceiling 1.675 TFLOPS; compute utilization 0.3% (the 989-TFLOPS ceiling is nowhere near).

Diagnosis: strongly memory-bound — more GPU compute won't help; bandwidth pressure must be reduced — apply Weight-Only Quantization and Mixed Precision (W4A16, turning 280GB → 70GB; AI stays the same but bandwidth pressure drops 4×) and increase batch size (multiple requests share the weight reads, so AI rises linearly with batch).

2. Prefill Phase ​

Llama-2-70B on a single H100 SXM5, single-request prefill with a 2048-token prompt:

  • Compute: ~2 × 2048 × 70B FLOPS ≈ 286 TFLOPS;
  • Bandwidth: ~140GB (weights read once; prompt input is negligible);
  • AI = 286 × 10¹² / 280 × 10⁹ ≈ 1020 FLOPS/Byte;
  • On the Roofline, AI = 1020 exceeds the ridge point of 295 → compute-bound;
  • Attainable compute = 989 TFLOPS (pinned against the compute ceiling).

Diagnosis: prefill is compute-bound, so adding compute (upgrading to B200) genuinely works. This is why vLLM and PagedAttention schedules prefill and decode separately — they have different bottlenecks and must be optimized separately.

3. Elementwise Operators ​

LayerNorm has AI ≈ 2.5, far below the ridge point of 295:

  • Attainable compute = 3.35TB/s × 2.5 ≈ 8.4 TFLOPS (0.8% of 989 TFLOPS);
  • If not fused, measured performance reads the input and output multiple times — fusing 3-4 elementwise operators into one kernel with Kernel Fusion and Custom Kernels leaves AI unchanged but cuts read/write passes from 4× to 1× — effective bandwidth ×4.

5. Choosing Optimization Strategies with Roofline ​

Roofline gives you not just a diagnosis but a priority order for optimization:

Operator PositionAIBottleneckOptimization Direction
Far left (AI ≈ 0.5)Extremely lowBandwidthQuantization to cut bytes, larger batch, kernel fusion
Mid-left (AI ~50)Low-mediumBandwidthKernel fusion, loop-order tuning, tiling
Near the ridge (AI ~295)CriticalBothOptimizable from either side — check how far measured performance sits from the Roofline
Just right of the ridge (AI ~500)Medium-highComputeReduce computation (distillation, pruning), use Tensor Cores
Far right (AI ~2000)HighComputeHigher-compute hardware, lower precision (FP8/FP4)

A General Optimization Mantra

6. Limitations of Roofline ​

Roofline is a simplified model; real operators are also affected by:

  1. Cache hits: Roofline assumes all data comes from HBM, but L2 hits make effective bandwidth much higher (even approaching SRAM bandwidth);
  2. Intrinsic operator parallelism: a matmul running on Tensor Cores vs. plain CUDA Cores differs by tens of times — Roofline doesn't distinguish;
  3. Instruction mix: an operator mixing add/sub/mul/div may not reach peak utilization (Tensor Cores only favor dense GEMM);
  4. Memory access pattern: coalesced vs. random access differs by 5-10× in bandwidth — Roofline counts totals but doesn't look at patterns;
  5. Latency hiding: GPUs hide HBM latency through multi-warp concurrency, but bandwidth still can't be saturated when there aren't enough warps.

Roofline Gives an Upper Bound, Not a Measurement

The Roofline gives a theoretical ceiling — if your operator measures at only 30% of its Roofline ceiling, the missing 70% comes from other factors (cache, warp occupancy, bank conflicts, instruction mix). To chase that, use Nsight Compute for single-operator profiling. See GPU Architecture and Optimization and Inference Benchmarking in Practice.

7. Trade-offs ​

  • Theory vs. measurement: Roofline is the starting point — it tells you "whether to optimize compute or bandwidth," but the actual gain must be measured;
  • Compute vs. bandwidth when choosing hardware: compute-bound tasks (training) favor high-compute cards (B200); memory-bound tasks (LLM inference) favor high-bandwidth cards (H200);
  • Quantization vs. distillation: decide by operator position — memory-bound favors quantization first (directly cuts bytes); compute-bound calls for distillation to cut computation;
  • Budget allocation: small teams should first reach the Roofline ceiling in software (vLLM + AWQ + kernel fusion) before considering hardware upgrades — in many cases software alone captures 80% of the gain.

Further Reading ​

References ​