← All parts of this equation

Equation 1 · Part 2 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol α

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)
α\alpha

What this part means

α is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

α is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.