← All parts of this equation

Equation 6 · Part 4 · Comparing the Main Approaches to Training Data and Synthetic Data

Symbol y

Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big)
yy

What this part means

y is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

y is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing ztz_t and zsz_s for teacher and student logits, σ\sigma for the softmax, y for the hard label, and T for a softening temperature, the canonical objective blends the two signals: Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big). The point of writing it out is the structure it exposes, not the arithmetic: the teacher term is fixed throughout training, so nothing the student produces ever feeds back…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.