Equation 7 · From n-Grams to Reasoning Models: A Technical History of the Language Model
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L
L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol N
N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol N_c
occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol alpha_N
alph is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: . Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…
Read the full surrounding passage
By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: . Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and the large models of the era were significantly undertrained. Chinchilla, at 70 billion parameters trained on roughly four times more data than the 280-billion-parameter Gopher at the same compute budget, outperformed it and reached 67.5% on MMLU [ 17 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to From n-Grams to Reasoning Models: A Technical History of the Language Model