A GPU has two resources: compute and memory bandwidth. Ops:byte describes the hardware. Arithmetic intensity describes the algorithm. Compare the two and you know which resource is your bottleneck — and therefore which optimizations are pointless.
The practical reading is the gap between your point and the horizontal roof — that's wasted silicon. The book's worked example lands at AI 62 on an H100, which means it can attain 62 × 3.35 TB/s ≈ 208 TFLOP/s out of 989 — about 21% of peak. Buying a faster-compute GPU would change nothing.
| Workload | Arithmetic intensity | Attainable | % of peak | Regime |
|---|
Hover a row or a point to link the two. The book's own worked example is the highlighted row.
| GPU / precision | Peak (dense) | Bandwidth | Ops:byte | Implication |
|---|---|---|---|---|
| A100 80GB SXM · FP16 | 312 TFLOPS | 2.04 TB/s | 153 | Lowest bar to clear — easiest to reach compute bound |
| H200 SXM · FP16 | 989 TFLOPS | 4.80 TB/s | 206 | Same compute as H100, more bandwidth → better for memory-bound decode |
| B200 · BF16 | 2,250 TFLOPS | 8.00 TB/s | 281 | Compute and bandwidth scaled together |
| H100 SXM · FP16 ← book's example | 989 TFLOPS | 3.35 TB/s | 295 | Needs 295 FLOPs per byte accessed to be balanced |
| L40S · FP16 | 362 TFLOPS | 0.86 TB/s | 419 | Bandwidth-starved → punishing for decode workloads |
| H100 SXM · FP8 | 1,979 TFLOPS | 3.35 TB/s | 591 | Doubling compute doubles the bar you must clear |
| B200 · FP4 | 9,000 TFLOPS | 8.00 TB/s | 1,125 | Low precision makes compute cheap, so memory dominates even harder |
A "better" GPU often has a higher ops:byte ratio, which makes it harder to keep busy. Quantizing weights to FP8 halves your bytes — but it also doubles peak compute, so the ridge point moves right just as fast as your arithmetic intensity does. The win from quantization is smaller weights and faster reads, not a free escape from the memory wall.
Standard, unoptimized attention. N = 4096, d = 128, FP16 (2 bytes/value). Q, K, V are N×d; the intermediates S and P are N×N.
A 4096×4096 FP16 matrix is ~32 MiB — about one high-resolution RAW photo, written out and read back in for no reason.
| Memory moved | 8Nd + 8N² = 138 MB |
| Compute | 4N²d ≈ 8.6 GFLOP |
| Arithmetic intensity | 8.6e9 / 1.38e8 ≈ 62 |
| H100 ops:byte | 295 |
| Verdict | 62 < 295 → memory bound |
Rearranging the same terms gives a closed form — and it has a ceiling:
No sequence length rescues it. The 4N² of intermediate traffic grows exactly as fast as the 4N²d of compute, so the ratio pins near 64 — always left of every ridge point in the table above. This is the entire motivation for FlashAttention: never materialize S and P, and memory traffic collapses to 8Nd, so AI becomes ≈ N/2 and rises with sequence length. That's why the book says FlashAttention is especially useful for compute-bound work like prefill and video generation.
The book's summary bottleneck map: prefill / KV-cache construction = compute bound · decode / token generation = memory bound · image and video generation = compute bound. And the reason to care: "If a certain operation is bottlenecked on memory bandwidth, no amount of compute optimization will make the system faster, and vice versa."