← Mathematical compendium

Published equation contexts

Na=NN_a = N

Why this formula appears here

Fix the vocabulary first. Let N be total parameters, NaN_a the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , NaN_a = N and M ∝\propto N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (3)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Na=NN_a = N

Equation 5 · Foundation Models

Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Fix the vocabulary first. Let N be total parameters, NaN_a the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , NaN_a = N and M ∝\propto N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.

Meanings in this article

  • NaN_a: the parameters actually used to process a given token.
  • NN: total parameters.
Equation guide → · Article →
Na=NN_a = N

Equation 12 · Foundation Models

Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Mixture-of-experts breaks the identity NaN_a = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .

Meanings in this article

Equation guide → · Article →