Equation 29 · Part 3 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol i
What this part means
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Its job in the formula
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Symbol i→Article meaning
The passage around this formula
Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: . with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count — the whole strategy in one line.
Learn the underlying idea
A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.
Open the illustrated functions: inputs become outputs guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.