← Mathematical compendium

Published equation contexts

M∝NM \propto N

Why this formula appears here

Fix the vocabulary first. Let N be total parameters, NaN_a the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , NaN_a = N and M ∝\propto N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.

Read the full article-specific guide →

Read the representative guide

MM

Symbol M

M is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

M∝NM \propto N

Equation 6 · Foundation Models

Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Fix the vocabulary first. Let N be total parameters, NaN_a the parameters actually used to process a given token, D training tokens, and M the memory that must be resident to serve. For the dense decoder-only transformer that remains the default architecture [ 1 ] , NaN_a = N and M ∝\propto N . Every method below breaks one of those two identities, and which one it breaks determines what it is good for.

Meanings in this article

  • NN: total parameters.
Equation guide → · Article →