← Mathematical compendium

Published equation contexts

σ(z)i=exp⁡(zi/T)∑jexp⁡(zj/T)\sigma(z)_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Why this formula appears here

Distillation recovers it by training the small “student” network against the large “teacher” network’s full output distribution rather than against the label alone. Write the teacher’s and student’s pre-softmax outputs for a given input as ztz_t and zsz_s . A softmax with a temperature T is σ(z)i=exp⁡(zi/T)∑jexp⁡(zj/T)\sigma(z)_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}. At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s…

Read the full article-specific guide →

Read the representative guide

jj

Symbol j

j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
zjz_j

Symbol z_j

zjz_j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
∑jexp⁡(zj/T)\sum_j \exp(z_j / T)

Denominator: sum_j exp(z_j / T)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →
jj

Starting index or lower bound: j

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

σ(z)i=exp⁡(zi/T)∑jexp⁡(zj/T)\sigma(z)_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Equation 4 · Edge AI & Electronics

How a Model Actually Gets Small Enough to Run on a Phone

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Distillation recovers it by training the small “student” network against the large “teacher” network’s full output distribution rather than against the label alone. Write the teacher’s and student’s pre-softmax outputs for a given input as ztz_t and zsz_s . A softmax with a temperature T is σ(z)i=exp⁡(zi/T)∑jexp⁡(zj/T)\sigma(z)_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}. At T=1 this is the ordinary softmax used for prediction. Raising T flattens the distribution, pulling the small probabilities assigned to wrong classes up toward visibility, which is exactly the dark knowledge the method wants to expose. The training objective blends two terms: ordinary cross-entropy against the true label, and a match between the student’s and teacher’s…

Equation guide → · Article →