Equation 29 · Part 17 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Starting index or lower bound: i in TopK(G(x))
What this part means
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Its job in the formula
i in TopK(G(x)) appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Starting index or lower bound: i in TopK(G(x))→Article meaning
The passage around this formula
Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: . with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count — the whole strategy in one line.
Learn the underlying idea
Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.
Open the illustrated sums and products: repeat an operation over an index guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.