← All parts of this equation

Equation 1 · Part 7 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol z_t

Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)
ztz_t

What this part means

the teacher’s and student’s output logits.

Its job in the formula

ztz_t is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label.

The passage around this formula

…a small “student” network to match a large “teacher” network’s full output distribution, not merely its top label — is usually written with a softened distribution and a temperature parameter: Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big). where ztz_t and zsz_s are the teacher’s and student’s output logits, σ\sigma a softmax, T a temperature that softens the distribution, and y the ground-truth label. The second term is the entire point of the method: it transfers the relative probability the teacher assigns to every wrong answer, not just…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.