Equation 12 · Part 1 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
Symbol N_a
What this part means
is part of the quantity the equation computes from the expression on the right.
Its job in the formula
is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol N_a→Article meaning
The passage around this formula
Mixture-of-experts breaks the identity = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [4] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ↗
- [5] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ↗
These citations provide research context; check each source for the exact claim it supports.