← Mathematical compendium

Published equation contexts

LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right)

Why this formula appears here

At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right). The T2T^2 factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/T2T^2 , multiplying the term by T2T^2 keeps the relative contribution of the two loss terms…

Read the full article-specific guide →

Read the representative guide

LKD\mathcal{L}_{\mathrm{KD}}

Symbol L_KD

LKL_KD is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
LCE\mathcal{L}_{\mathrm{CE}}

Symbol L_CE

LCL_CE is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right)

Equation 7 · Edge AI & Electronics

How a Model Actually Gets Small Enough to Run on a Phone

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right). The T2T^2 factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/T2T^2 , multiplying the term by T2T^2 keeps the relative contribution of the two loss terms…

Meanings in this article

  • T2T^2: the square of T; tuned as a hyperparameter [ 1 ].
  • TT: tuned as a hyperparameter [ 1 ].
Equation guide → · Article →