Equation 18 · Part 1 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol L
What this part means
L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol L→Article meaning
The passage around this formula
with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The data-efficient small-model strategy asks a different question with the same law: given a fixed, small N set by the deployment target, what choice of D — and, crucially, what quality of D — minimizes L ? Fixing the small side of the equation first and spending the freed budget on data rather than on parameters is the strategy’s entire premise.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [4] Training Compute-Optimal Large Language Models ↗
- [5] Textbooks Are All You Need ↗
- [6] TinyStories: How Small Can Language Models Be and Still Speak Coherent English? ↗
These citations provide research context; check each source for the exact claim it supports.