Equation 29 · Part 7 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol k
What this part means
k occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Its job in the formula
k occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Full expression→Symbol k→Article meaning
The passage around this formula
Shazeer and colleagues stated the underlying argument for conditional computation directly: “a trainable gating network determines a sparse combination of experts to use for each example,” a mechanism they showed could scale model capacity by “over 1000x” while keeping the compute spent on any one example roughly constant [ 11 ] . The now-standard form of a sparse mixture-of-experts layer routes each token to a small top- k subset of E available experts: . with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count — the whole strategy in one line.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.