← Mathematical compendium

Published equation contexts

qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Why this formula appears here

Hinton, Vinyals and Dean’s 2015 paper, “Distilling the Knowledge in a Neural Network,” is the one the field actually cites, and it earned that position by generalizing the 2006 idea and giving it a mechanism simple enough to fit into any neural network’s training pipeline. Instead of training a small model against a large model’s hard output labels, distillation trains it against the large “teacher” model’s full, softened probability distribution over classes — the softmax output computed at a raised temperature T : qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}. where ziz_i are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is…

Read the full article-specific guide →

Read the representative guide

jj

Symbol j

j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
zjz_j

Symbol z_j

zjz_j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
∑jexp⁡(zj/T)\sum_j \exp(z_j / T)

Denominator: sum_j exp(z_j / T)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →
jj

Starting index or lower bound: j

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Equation 10 · Edge AI & Electronics

A History of Small and On-Device AI

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Hinton, Vinyals and Dean’s 2015 paper, “Distilling the Knowledge in a Neural Network,” is the one the field actually cites, and it earned that position by generalizing the 2006 idea and giving it a mechanism simple enough to fit into any neural network’s training pipeline. Instead of training a small model against a large model’s hard output labels, distillation trains it against the large “teacher” model’s full, softened probability distribution over classes — the softmax output computed at a raised temperature T : qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}. where ziz_i are the teacher’s pre-softmax logits. “Using a higher value for T produces a softer probability distribution over classes” [ 4 ] , and it is…

Meanings in this article

Equation guide → · Article →