← Mathematical compendium

Published equation contexts

y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})

Why this formula appears here

Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}}). with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count NtotalN_{\text{total}} — the whole strategy in one line.

Read the full article-specific guide →

Read the representative guide

xx

Symbol x

x appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
GG

Symbol G

G appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
EE

Symbol E

E occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
CtokdenseC_{\text{tok}}^{\text{dense}}

Symbol C_tok^dense

CtC_tokdk^dense is one factor in the product that computes the quantity on the left.

Read this term in its guide →
NtotalN_{\text{total}}

Symbol N_total

NtN_total is one factor in the product that computes the quantity on the left.

Read this term in its guide →
i∈TopK(G(x))i \in \mathrm{TopK}(G(x))

Starting index or lower bound: i in TopK(G(x))

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})

Equation 29 · Edge AI & Electronics

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

This equation gives an approximation: it relates the quantities while allowing an approximation.

Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}}). with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count NtotalN_{\text{total}} — the whole strategy in one line.

Equation guide → · Article →