Symbol L_KD
D is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: . The factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/ , multiplying the term by keeps the relative contribution of the two loss terms…
D is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →α is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →E is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →y is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →σ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 7 · Edge AI & Electronics
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s temperature-softened distributions, measured by KL divergence: . The factor is not decorative. Hinton and colleagues note that because the magnitude of the gradients produced by the soft-target term scales as 1/ , multiplying the term by keeps the relative contribution of the two loss terms…