← All parts of this equation

Equation 12 · Part 2 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

Symbol N

Na=NN_a = N
NN

What this part means

the not.

Its job in the formula

N is part of the quantity the equation computes from the expression on the right.

Where the article explains it

Training compute scales with NaN_a , not N , so a sparse model can hold far more knowledge for the same training FLOPs.

The passage around this formula

Mixture-of-experts breaks the identity NaN_a = N . Shazeer and colleagues introduced the sparsely-gated MoE layer, in which a learned gate selects a small subset of expert sub-networks per example, allowing parameter counts far beyond what could be densely activated at the same compute [ 4 ] . Fedus, Zoph, and Shazeer simplified it decisively: the Switch layer routes each token to exactly one expert rather than the top- k , which reduced routing computation and communication cost while preserving quality, and they scaled the approach to trillion-parameter models [ 5 ] .

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.