Can the Grace CPU access the GPU's HBM? Yes, over NVLink-C2C

On a Grace Blackwell superchip the "unified memory" is not a third pool. It is the GPU's HBM3e plus the CPU's LPDDR5X mapped into one coherent address space. The CPU reaches HBM, and the GPU reaches LPDDR5X, with ordinary loads and stores. What differs is the path each access takes and how much bandwidth that path has.

One GB200 superchip: two physical memories, one address space

The highlighted route is a CPU load or store that lands in HBM. It crosses the chip-to-chip link, so it is limited to the link's bandwidth, not HBM's. Hover or focus any part for details.

Capacities and bandwidths are NVIDIA's rated figures for GB200: 192 GB HBM3e at 8 TB/s per GPU, up to 480 GB LPDDR5X at about 500 GB/s per Grace, NVLink-C2C at 900 GB/s bidirectional per GPU, NVLink 5 at 1.8 TB/s bidirectional per GPU.

Bandwidth to a byte depends on who is asking and where it lives

Links are shown per direction (half of NVIDIA's bidirectional figure). Memory bandwidths are totals. "As if one huge pool" holds for correctness; for performance the CPU-to-HBM route has about one eighteenth of the GPU's local bandwidth.

Latency moves the same way: a CPU access into HBM, or a GPU access into LPDDR5X, is noticeably slower than a local access. The driver's page migration and cudaMemPrefetchAsync / cudaMemAdvise exist to keep hot data on the side that uses it.

What "unified" means here

  • One virtual address space. A pointer is valid on both sides. For memory from malloc or mmap the GPU uses the CPU's page tables through Address Translation Services (ATS), so nothing has to be registered or copied first.
  • Hardware coherence over NVLink-C2C. The CPU can cache lines that live in HBM, and the GPU can cache lines that live in LPDDR5X, with the hardware keeping both views consistent. System-scope atomics work across the two.
  • Two physical pools, visible to the OS. Linux exposes each GPU's HBM as a CPU-less NUMA node, so even numactl --membind can place ordinary allocations in HBM, and the CPU can touch them.

What the CPU can and cannot touch

  • System-allocated and managed memory, wherever its pages live. If a page of malloc or cudaMallocManaged memory is resident in HBM, a CPU load reads it over C2C. The driver may also migrate hot pages between HBM and LPDDR5X.
  • Not cudaMalloc memory. That stays GPU-private in the CUDA programming model. The CPU cannot dereference such a pointer, even though the bytes physically sit in the same HBM.
  • Extended GPU Memory (EGM) is the GPU-side name. It lets a GPU address the LPDDR5X of its own Grace and, over NVLink, the LPDDR5X of other superchips in the rack, so a GPU's reachable memory grows toward the rack's roughly 30 TB. The book uses EGM as a synonym for unified memory; strictly it describes the GPU extending into CPU memory, not the CPU reaching HBM.

Figure for chapter 2, "AI System Hardware Overview". Colors follow the OS light or dark setting.