← All parts of this equation

Equation 29 · Part 8 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

Symbol E

y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})
EE

What this part means

E occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

E occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}}). with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count NtotalN_{\text{total}} — the whole strategy in one line.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.