← All parts of this equation

Equation 1 · Part 4 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol σ

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)
σ\sigma

What this part means

the softmax.

Its job in the formula

σ is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label.

The passage around this formula

…full output distribution, not merely its top label — is usually written with a softened distribution and a temperature parameter: Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big). where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label. The second term is the entire point of the method: it transfers the relative probability the teacher assigns to every wrong answer, not just which answer was right, a far richer training signal per…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.