← Back to article

Equation 7 · From n-Grams to Reasoning Models: A Technical History of the Language Model

What does this equation mean?

L(N)≈(NcN)αN.L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

LL

Symbol L

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

NcN_c

Symbol N_c

NcN_c occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

αN\alpha_N

Symbol alpha_N

alphaNa_N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}. Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…
Read the full surrounding passage
By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}. Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and the large models of the era were significantly undertrained. Chinchilla, at 70 billion parameters trained on roughly four times more data than the 280-billion-parameter Gopher at the same compute budget, outperformed it and reached 67.5% on MMLU [ 17 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to From n-Grams to Reasoning Models: A Technical History of the Language Model

See this formula across 1 published context →

Browse the mathematical compendium →