Equation 8 · Part 5 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol A
What this part means
A occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Its job in the formula
A occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Full expression→Symbol A→Article meaning
The passage around this formula
…how large models are trained, applied in the opposite direction. Hoffmann and colleagues showed that contemporary large language models had been trained on too little data relative to their parameter count, and that for a fixed training budget, model size and training tokens should grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as . with N parameters and D training…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.