Symbol L_distill
istill is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing and for teacher and student logits, for the softmax, y for the hard label, and T for a softening temperature, the canonical…
istill is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →λ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →E is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →y is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →σ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →L is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →T is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 6 · Data & Training
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing and for teacher and student logits, for the softmax, y for the hard label, and T for a softening temperature, the canonical…