Equation 31 · Part 1 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol k
What this part means
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol k→Article meaning
The passage around this formula
with G(x) a learned gating distribution over experts. Compute per token scales with the active fraction k/E , not with the total parameter count — the whole strategy in one line.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [11] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ↗
- [12] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ↗
- [13] Mixtral of Experts ↗
These citations provide research context; check each source for the exact claim it supports.