← Back to article

Equation 12 · A History of Small and On-Device AI

What does this equation mean?

TT

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

TT

Symbol T

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where ziz_i are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is precisely that softness — the relative probabilities the teacher assigns to the wrong answers, not only the single correct label — that carries information a hard label discards: how confidently a teacher model preferred one wrong answer over another says far more about the shape of its decision boundary than one correct label does. The paper states the mechanism directly: “Knowledge is transferred to the distilled model by training it on a transfer set and using a soft target distribution for each case in the transfer set that is produced by…
Read the full surrounding passage
where ziz_i are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is precisely that softness — the relative probabilities the teacher assigns to the wrong answers, not only the single correct label — that carries information a hard label discards: how confidently a teacher model preferred one wrong answer over another says far more about the shape of its decision boundary than one correct label does. The paper states the mechanism directly: “Knowledge is transferred to the distilled model by training it on a transfer set and using a soft target distribution for each case in the transfer set that is produced by using the cumbersome model with a high temperature in its softmax” [ 4 ] . On a speech-recognition acoustic model, the authors reported that “more than 80% of the improvement in frame classification accuracy achieved by using an ensemble of 10 models is transferred to the distilled model” [ 4 ] — most of what ten expensive models knew, recovered in a single model built to run where the ensemble could not.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to A History of Small and On-Device AI

Browse the mathematical compendium →