C ≈ 6ND: how the 100-trillion-parameter training estimate is built

Chapter 1 says a dense 100-trillion-parameter model trained on about 29 trillion tokens needs about 1.74 × 1028 FLOPs. The formula behind that number counts how many times each parameter is touched per training token.

The 6 is FLOPs per parameter per token

Dense transformer. Every parameter is used once per token in the forward pass and twice more in the backward pass. One use = one multiply + one add = 2 FLOPs.

Why 2 FLOPs per parameter

For one token, a layer computes yi = Σj Wij xj. Each weight Wij is touched exactly once: one multiply (Wijxj) and one add (accumulate into yi). So a matrix with P entries costs 2P FLOPs per token, whatever its shape.

Backprop repeats this twice: once to push the gradient to the input (W⊤ has the same P entries) and once to form the weight gradient (an outer product with P outputs). Total: 2P + 2P + 2P = 6P.

Ignored by the heuristic: attention scores (QK⊤, softmax, the multiply by V), layer norms, activations, embeddings. These add a modest fraction unless the context length is very long relative to the hidden size.

Plugging in the 100T model

N = 1 × 1014 parameters (100 trillion)
D = 2.9 × 1013 tokens (29 trillion)

C= 6 × N × D
= 6 × (1 × 1014) × (2.9 × 1013)
= (6 × 2.9) × 1014 + 13
= 17.4 × 1027
≈ 1.74 × 1028 FLOPs

Read 1.74 × 1028 as 17.4 billion billion billion floating-point operations. Doubling the tokens doubles C. A sparse MoE model uses active parameters for N.

Training compute by the 6ND heuristic

The 100-trillion-parameter model needs about 460× the compute of Llama 3.1 405B and 55,000× GPT-3. Published figures match the heuristic: GPT-3 reported 3.14 × 1023, Llama 3.1 reported 3.8 × 1025.

Sources: N and D from each model's paper. MoE (mixture-of-experts) models count the parameters active per token, not the total. Hover or focus a bar for details.

Figures for chapter 1, "Toward 100-Trillion-Parameter Models". Colors follow the OS light or dark setting.