← All parts of this equation

Equation 6 · Part 14 · Comparing the Main Approaches to Training Data and Synthetic Data

superscript

Ldistill=(1−λ) LCE(y, σ(zs))+λ T2 DKL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{distill}} = (1-\lambda)\,\mathcal{L}_{\mathrm{CE}}(y,\, \sigma(z_s)) + \lambda\, T^2 \, D_{\mathrm{KL}}\big(\sigma(z_t/T) \,\|\, \sigma(z_s/T)\big)
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [ 1 ] . Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing ztz_t and zsz_s for teacher and student logits, σ\sigma for the softmax, y for the hard label, and T for a softening temperature, the canonical…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.