← All parts of this equation

Equation 8 · Part 1 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol L

L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
LL

What this part means

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

…also be doubled” [ 4 ] . The underlying loss law is commonly written as L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}. with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The data-efficient small-model strategy asks a different question with the same law: given a fixed, small N set by the deployment target, what choice of D — and, crucially, what quality of D — minimizes L ? Fixing the small side of the equation first and spending the freed budget on data rather…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.