Equation 12 · Part 2 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
Symbol N
What this part means
the not.
Its job in the formula
N is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol N→Article meaning
Where the article explains it
Training compute scales with , not N , so a sparse model can hold far more knowledge for the same training FLOPs.
The passage around this formula
Mixture-of-experts breaks the identity = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [4] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ↗
- [5] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ↗
These citations provide research context; check each source for the exact claim it supports.