Equation 1 · Part 6 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol T^2
What this part means
the square of T; the temperature that softens the distribution.
Its job in the formula
is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol T^2→Article meaning
Where the article explains it
where and are the teacher’s and student’s output logits, a softmax, T a temperature that softens the distribution, and y the ground-truth label.
The passage around this formula
The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.