Equation 1 · Part 11 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
superscript
superscript
What this part means
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
Its job in the formula
A raised mark can be a power or an index. Its position and the surrounding notation determine which.
Full expression→superscript→Article meaning
The passage around this formula
The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.