Equation 8 · Part 3 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol D
What this part means
D occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Its job in the formula
D occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Symbol D→Article meaning
The passage around this formula
…grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as . with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The data-efficient small-model strategy asks a different question with the same law: given a fixed, small N set by the deployment target, what choice of D — and, crucially, what quality of D —…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.