← Mathematical compendium

Published equation contexts

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)

Why this formula appears here

The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…

Read the full article-specific guide →

Read the representative guide

Ldistill\mathcal{L}_{\text{distill}}

Symbol L_distill

LdL_distill is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)

Equation 1 · Edge AI & Electronics

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The oldest of the four strategies starts from an asset that is already paid for: a large model has already been trained, at whatever cost that took, and the cost is sunk. Compression treats the sunk cost as something to copy rather than repeat. Hinton, Vinyals, and Dean set out the core argument in 2015, motivated by a practical deployment problem with large ensembles: “making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users” [ 1 ] . Their proposed fix — train a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written…

Meanings in this article

  • yy: the ground-truth label.
  • σ\sigma: the softmax.
  • zsz_s: the teacher’s and student’s output logits.
  • T2T^2: the square of T; the temperature that softens the distribution.
  • ztz_t: the teacher’s and student’s output logits.
  • TT: the temperature that softens the distribution.
Equation guide → · Article →