← Mathematical compendium

Published equation contexts

Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big)

Why this formula appears here

Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing ztz_t and zsz_s for teacher and student logits, σ\sigma for the softmax, y for the hard label, and T for a softening temperature, the canonical…

Read the full article-specific guide →

Read the representative guide

Ldistill\mathcal{L}_{\mathrm{distill}}

Symbol L_distill

LdL_distill is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
LCE\mathcal{L}_{\mathrm{CE}}

Symbol L_CE

LCL_CE is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
DKLD_{\mathrm{KL}}

Symbol D_KL

DKD_KL is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big)

Equation 6 · Data & Training

Comparing the Main Approaches to Training Data and Synthetic Data

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing ztz_t and zsz_s for teacher and student logits, σ\sigma for the softmax, y for the hard label, and T for a softening temperature, the canonical…

Meanings in this article

Equation guide → · Article →