Equation 12 · Part 4 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
Mixture-of-experts breaks the identity = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the surrounding passage
- [4] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ↗
- [5] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ↗
These citations provide research context; check each source for the exact claim it supports.