← All parts of this equation

Equation 1 · Part 6 · Comparing the Main Approaches to AI Inference Economics

≈

CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

For inference cost specifically, what matters is that per-token compute tracks activated parameters, not total parameters. Approximating compute per token as proportional to the parameter count actually touched [ 13 , 11 ] , the ratio of dense to MoE compute at equal activated size is approximately CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}. for a dense model with NdenseN_{\mathrm{dense}} parameters compared against an MoE model activating NactiveN_{\mathrm{active}} of its NtotalN_{\mathrm{total}} parameters per token. What this ratio hides is exactly what a compute-only comparison always hides: memory. Serving an MoE model requires holding all NtotalN_{\mathrm{total}} parameters resident — on one device or, more often, sharded across…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.