← All parts of this equation

Equation 15 · Part 1 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

Symbol N

NN
NN

What this part means

the not.

Its job in the formula

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

Training compute scales with NaN_a , not N , so a sparse model can hold far more knowledge for the same training FLOPs.

The passage around this formula

The economics are attractive and frequently overstated. Training compute scales with NaN_a , not N , so a sparse model can hold far more knowledge for the same training FLOPs. Patterson and colleagues, computing energy and carbon for several large models including Switch Transformer and GPT-3, found that large but sparsely activated networks can consume less than one tenth the energy of large dense networks without sacrificing accuracy [ 10 ] .

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.