← All parts of this equation

Equation 29 · Part 16 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

superscript

y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}}). with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count NtotalN_{\text{total}} — the whole strategy in one line.

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.