← Back to article

Equation 7 · How a Model Actually Gets Small Enough to Run on a Phone

What does this equation mean?

LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

LKD\mathcal{L}_{\mathrm{KD}}

Symbol L_KD

LKL_KD is part of the quantity the equation computes from the expression on the right.

Understand this part →

α\alpha

Symbol α

α is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

LCE\mathcal{L}_{\mathrm{CE}}

Symbol L_CE

LCL_CE is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

yy

Symbol y

y is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

σ\sigma

Symbol σ

σ is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

zsz_s

Symbol z_s

zsz_s is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

T2T^2

Symbol T^2

the square of T; tuned as a hyperparameter [ 1 ].

Understand this part →

ztz_t

Symbol z_t

ztz_t is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

TT

Symbol T

tuned as a hyperparameter [ 1 ].

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right). The T2T^2 factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/T2T^2 , multiplying the term by T2T^2 keeps the relative contribution of the two loss terms…
Read the full surrounding passage
At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right). The T2T^2 factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/T2T^2 , multiplying the term by T2T^2 keeps the relative contribution of the two loss terms roughly stable as T is tuned as a hyperparameter [ 1 ] . Without it, raising the temperature to expose more dark knowledge would simultaneously and silently shrink the gradient signal doing the exposing.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How a Model Actually Gets Small Enough to Run on a Phone

See this formula across 1 published context →

Browse the mathematical compendium →