← Mathematical compendium

Published equation contexts

L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}

Why this formula appears here

By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}. Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…

Read the full article-specific guide →

Read the representative guide

LL

Symbol L

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
αN\alpha_N

Symbol alpha_N

alphaNa_N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

L(N)≈(NcN)αN.L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}.

Equation 7 · Foundation Models

From n-Grams to Reasoning Models: A Technical History of the Language Model

This equation gives an approximation: it relates the quantities while allowing an approximation.

By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}. Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…

Equation guide → · Article →