← All parts of this equation

Equation 10 · Part 1 · A History of Small and On-Device AI

Symbol q_i

qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}
qiq_i

What this part means

qiq_i is part of the quantity the equation computes from the expression on the right.

Its job in the formula

qiq_i is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Hinton, Vinyals and Dean’s 2015 paper, “Distilling the Knowledge in a Neural Network,” is the one the field actually cites, and it earned that position by generalizing the 2006 idea and giving it a mechanism simple enough to fit into any neural network’s training pipeline. Instead of training a small model against a large model’s hard output labels, distillation trains it against the large “teacher” model’s full, softened probability distribution over classes — the softmax output computed at a raised temperature T : qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}. where ziz_i are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.