Equation 7 · How a Model Actually Gets Small Enough to Run on a Phone
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L_KD
D is part of the quantity the equation computes from the expression on the right.
Symbol α
α is one of the signed contributions combined to compute the quantity on the left.
Symbol L_CE
E is one of the signed contributions combined to compute the quantity on the left.
Symbol y
y is one of the signed contributions combined to compute the quantity on the left.
Symbol σ
σ is one of the signed contributions combined to compute the quantity on the left.
Symbol z_s
is one of the signed contributions combined to compute the quantity on the left.
Symbol z_t
is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: . The factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/ , multiplying the term by keeps the relative contribution of the two loss terms…
Read the full surrounding passage
At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: . The factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/ , multiplying the term by keeps the relative contribution of the two loss terms roughly stable as T is tuned as a hyperparameter [ 1 ] . Without it, raising the temperature to expose more dark knowledge would simultaneously and silently shrink the gradient signal doing the exposing.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How a Model Actually Gets Small Enough to Run on a Phone