← All parts of this equation

Equation 1 · Part 9 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

=

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.