Symbol L
L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: . Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…
L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →alph is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 7 · Foundation Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: . Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…
Equation guide → · Article →