← Mathematical compendium

Published equation contexts

CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}

Why this formula appears here

For inference cost specifically, what matters is that per-token compute tracks activated parameters, not total parameters. Approximating compute per token as proportional to the parameter count actually touched [ 13 , 11 ] , the ratio of dense to MoE compute at equal activated size is approximately CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}. for a dense model with NdenseN_{\mathrm{dense}} parameters compared against an MoE model activating NactiveN_{\mathrm{active}} of its NtotalN_{\mathrm{total}} parameters per token. What this ratio hides is exactly what a compute-only comparison always hides: memory. Serving an MoE model requires holding all NtotalN_{\mathrm{total}} parameters resident — on one device or, more often, sharded across…

Read the full article-specific guide →

Read the representative guide

CdenseC_{\mathrm{dense}}

Symbol C_dense

CdC_dense occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
CMoEC_{\mathrm{MoE}}

Symbol C_MoE

CMC_MoE occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
NdenseN_{\mathrm{dense}}

Symbol N_dense

NdN_dense occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
NactiveN_{\mathrm{active}}

Symbol N_active

NaN_active occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}

Equation 1 · Inference Economics

Comparing the Main Approaches to AI Inference Economics

This equation gives an approximation: it relates the quantities while allowing an approximation.

For inference cost specifically, what matters is that per-token compute tracks activated parameters, not total parameters. Approximating compute per token as proportional to the parameter count actually touched [ 13 , 11 ] , the ratio of dense to MoE compute at equal activated size is approximately CdenseCMoE≈NdenseNactive\frac{C_{\mathrm{dense}}}{C_{\mathrm{MoE}}} \approx \frac{N_{\mathrm{dense}}}{N_{\mathrm{active}}}. for a dense model with NdenseN_{\mathrm{dense}} parameters compared against an MoE model activating NactiveN_{\mathrm{active}} of its NtotalN_{\mathrm{total}} parameters per token. What this ratio hides is exactly what a compute-only comparison always hides: memory. Serving an MoE model requires holding all NtotalN_{\mathrm{total}} parameters resident — on one device or, more often, sharded across…

Equation guide → · Article →