← Mathematical compendium

Published equation contexts

L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}

Why this formula appears here

Its intellectual foundation is the same scaling-law literature that shaped how large models are trained, applied in the opposite direction. Hoffmann and colleagues showed that contemporary large language models had been trained on too little data relative to their parameter count, and that for a fixed training budget, model size and training tokens should grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}. with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The…

Read the full article-specific guide →

Read the representative guide

LL

Symbol L

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
DD

Symbol D

D occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
EE

Symbol E

E is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
NαN^{\alpha}

Symbol N^α

N^α occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
DβD^{\beta}

Symbol D^β

D^β occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}

Equation 8 · Edge AI & Electronics

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

This equation gives an approximation: it relates the quantities while allowing an approximation.

Its intellectual foundation is the same scaling-law literature that shaped how large models are trained, applied in the opposite direction. Hoffmann and colleagues showed that contemporary large language models had been trained on too little data relative to their parameter count, and that for a fixed training budget, model size and training tokens should grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}. with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The…

Meanings in this article

Equation guide → · Article →