← All parts of this equation

Equation 1 · Part 3 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol y

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)
yy

What this part means

the ground-truth label.

Its job in the formula

y is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label.

The passage around this formula

…written with a softened distribution and a temperature parameter: Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big). where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label. The second term is the entire point of the method: it transfers the relative probability the teacher assigns to every wrong answer, not just which answer was right, a far richer training signal per example than a raw label alone provides. That is the rationale, stated…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.