← All parts of this equation

Equation 6 · Part 1 · Comparing the Main Approaches to Training Data and Synthetic Data

Symbol L_distill

Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big)
Ldistill\mathcal{L}_{\mathrm{distill}}

What this part means

LdL_distill is part of the quantity the equation computes from the expression on the right.

Its job in the formula

LdL_distill is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing ztz_t and zsz_s for teacher and student logits, σ\sigma for the softmax, y for the hard label, and T for a softening temperature, the canonical…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.