← Back to article

Equation 16 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

What does this equation mean?

MM

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

MM

Symbol M

M is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

What sparsity does not break is M . Every expert must be resident somewhere at serving time even though any given token touches one. Since decoding is bound by memory traffic rather than arithmetic — Gholami and colleagues report peak server FLOPS scaling at roughly 3.0× every two years against DRAM and interconnect bandwidth at about 1.6× and 1.4× [ 9 ] — a sparse model’s advantage is real in training and much more conditional in serving. It buys quality per training FLOP and per activated parameter; it does not buy memory.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

Browse the mathematical compendium →