Published equation contexts
Why this formula appears here
Fix the vocabulary first. Let N be total parameters, the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , = N and M N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.
Read the representative guide
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (3)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · Foundation Models
Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Fix the vocabulary first. Let N be total parameters, the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , = N and M N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.
Meanings in this article
Equation guide → · Article →Equation 12 · Foundation Models
Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Mixture-of-experts breaks the identity = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .
Meanings in this article
Equation guide → · Article →Equation 17 · Foundation Models
Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
What sparse routing breaks: = N . Best when training compute is the binding constraint and memory is not.