← All parts of this equation

Equation 7 · Part 12 · How a Model Actually Gets Small Enough to Run on a Phone

subscript

LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right)
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right). The T2T^2 factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/T2T^2 , multiplying the term by T2T^2 keeps the relative contribution of the two loss terms…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.