Where the GB200 NVL72's "1.44 exaFLOPS" comes from

The chapter quotes 1.44 exaFLOPS for FP4 and 720 petaFLOPS for FP8. Both are NVIDIA datasheet peaks: one Blackwell GPU's rated tensor-core throughput, doubled by 2:1 structured sparsity, multiplied by the 72 GPUs in the rack.

GB200 NVL72 peak tensor throughput by precision

Each halving of the number format doubles the rate. The sparsity figure is the dense figure doubled again. The book's 1.44 EFLOPS and 720 PFLOPS are the two top sparse totals.

Source: NVIDIA GB200 NVL72 datasheet, whose footnote reads "All petaFLOPS and petaOPS are with Sparsity except FP64 which is dense." Per-GPU values are the datasheet totals divided by 72. Hover or focus a row for the per-GPU numbers.

The multiplication in the book

per GPU20 PFLOPS FP4, with sparsity
× GPUs18 compute trays × 2 superchips × 2 GPUs = 72
rack72 × 20 = 1,440 PFLOPS = 1.44 EFLOPS
per GPU10 PFLOPS FP8, with sparsity
rack72 × 10 = 720 PFLOPS

Without sparsity the same rack is 720 PFLOPS FP4 and 360 PFLOPS FP8. Dense training runs at the dense figure, and typically sustains 30 to 50 percent of it.

Where a GPU's peak itself comes from

A peak number is a count of tensor-core multiply-adds per clock, times the clock, times 2 FLOPs per multiply-add:

FLOPS = SMs × FMA per SM per clock × 2 × clock

NVIDIA published these inputs for Hopper, so the check is exact: 132 SMs × 2,048 dense FP16 FMA × 2 × 1.83 GHz = 989 TFLOPS, the H100 SXM datasheet value. For Blackwell NVIDIA has not published the SM count and tensor-core clock rates in the same detail, so the 20 PFLOPS per GPU is taken from its rated spec.

Two effects stack on top of that base rate:

  1. Narrower formats double throughput. The same tensor-core datapath processes twice as many 8-bit values as 16-bit ones per cycle, and twice as many 4-bit values again.
  2. 2:1 structured sparsity doubles it once more. When 2 of every 4 consecutive weights are zero, the hardware skips them and does only the nonzero half of the work. This applies only to weights pruned to that pattern.

The memory figures in the same passage

HBM3e72 GPUs × 192 GB = 13,824 GB = 13.5 TiB (the book's "13.5 TB", dividing by 1,024)
Grace LPDDR5X36 CPUs × 480 GB = 17,280 GB ≈ 16.9 TiB (datasheet: "up to 17 TB")
fast memory13.5 + 16.9 ≈ 30.4 TiB (datasheet: "up to 30 TB")

"Fast memory" is NVIDIA's name for the HBM plus CPU memory reachable over NVLink-C2C in one rack. It is not a single flat memory: remote pages have different performance, as the book notes.

Figure for chapter 1, section "NVIDIA's 'AI Supercomputer in a Rack'". Colors follow the OS light or dark setting.