Equation 10 · Part 4 · A History of Small and On-Device AI
Symbol j
What this part means
j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Its job in the formula
j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Symbol j→Article meaning
The passage around this formula
Hinton, Vinyals and Dean’s 2015 paper, “Distilling the Knowledge in a Neural Network,” is the one the field actually cites, and it earned that position by generalizing the 2006 idea and giving it a mechanism simple enough to fit into any neural network’s training pipeline. Instead of training a small model against a large model’s hard output labels, distillation trains it against the large “teacher” model’s full, softened probability distribution over classes — the softmax output computed at a raised temperature T : . where are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.