← Back to article

Equation 29 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

What does this equation mean?

y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withk
Divide byE
This relates toy(x)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

yy

Symbol y

y is part of the quantity the equation computes from the expression on the right.

Understand this part →

xx

Symbol x

x appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

GG

Symbol G

G appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

EiE_i

Symbol E_i

EiE_i is one factor in the product that computes the quantity on the left.

Understand this part →

CtokC_{\text{tok}}

Symbol C_tok

CtC_tok is one factor in the product that computes the quantity on the left.

Understand this part →

kk

Symbol k

k occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

EE

Symbol E

E occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

CtokdenseC_{\text{tok}}^{\text{dense}}

Symbol C_tok^dense

CtC_tokdk^dense is one factor in the product that computes the quantity on the left.

Understand this part →

NtotalN_{\text{total}}

Symbol N_total

NtN_total is one factor in the product that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
i∈TopK(G(x))i \in \mathrm{TopK}(G(x))

Starting index or lower bound: i in TopK(G(x))

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}}). with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count NtotalN_{\text{total}} — the whole strategy in one line.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

See this formula across 1 published context →

Browse the mathematical compendium →