Equation 8 · Part 12 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
superscript
superscript
What this part means
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
Its job in the formula
A raised mark can be a power or an index. Its position and the surrounding notation determine which.
Full expression→superscript→Article meaning
The passage around this formula
Its intellectual foundation is the same scaling-law literature that shaped how large models are trained, applied in the opposite direction. Hoffmann and colleagues showed that contemporary large language models had been trained on too little data relative to their parameter count, and that for a fixed training budget, model size and training tokens should grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as . with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.