C ≈ 6ND: how the 100-trillion-parameter training estimate is built
Chapter 1 says a dense 100-trillion-parameter model trained on about 29 trillion tokens needs about 1.74 × 1028 FLOPs. The formula behind that number counts how many times each parameter is touched per training token.
The 6 is FLOPs per parameter per token
Dense transformer. Every parameter is used once per token in the forward pass and twice more in the backward pass. One use = one multiply + one add = 2 FLOPs.
Why 2 FLOPs per parameter
For one token, a layer computes yi = Σj Wij xj. Each weight Wij is touched exactly once: one multiply (Wijxj) and one add (accumulate into yi). So a matrix with P entries costs 2P FLOPs per token, whatever its shape.
Backprop repeats this twice: once to push the gradient to the input (W⊤ has the same P entries) and once to form the weight gradient (an outer product with P outputs). Total: 2P + 2P + 2P = 6P.
Ignored by the heuristic: attention scores (QK⊤, softmax, the multiply by V), layer norms, activations, embeddings. These add a modest fraction unless the context length is very long relative to the hidden size.
Plugging in the 100T model
N = 1 × 1014 parameters (100 trillion)
D = 2.9 × 1013 tokens (29 trillion)
Read 1.74 × 1028 as 17.4 billion billion billion floating-point operations. Doubling the tokens doubles C. A sparse MoE model uses active parameters for N.
Training compute by the 6ND heuristic
The 100-trillion-parameter model needs about 460× the compute of Llama 3.1 405B and 55,000× GPT-3. Published figures match the heuristic: GPT-3 reported 3.14 × 1023, Llama 3.1 reported 3.8 × 1025.
| Model | Parameters (N) | Tokens (D) | C = 6ND (FLOPs) | Relative to 100T model |
|---|
Figures for chapter 1, "Toward 100-Trillion-Parameter Models". Colors follow the OS light or dark setting.