Equation 17 · Part 2 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
Symbol N
What this part means
the not.
Its job in the formula
N is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol N→Article meaning
Where the article explains it
Training compute scales with , not N , so a sparse model can hold far more knowledge for the same training FLOPs.
The passage around this formula
What sparse routing breaks: = N . Best when training compute is the binding constraint and memory is not.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [4] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer ↗
- [5] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity ↗
- [10] Carbon Emissions and Large Neural Network Training ↗
- [9] AI and Memory Wall ↗
These citations provide research context; check each source for the exact claim it supports.