Symbol C_dense
ense occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →Published equation contexts
For inference cost specifically, what matters is that per-token compute tracks activated parameters, not total parameters. Approximating compute per token as proportional to the parameter count actually touched [ 13 , 11 ] , the ratio of dense to MoE compute at equal activated size is approximately . for a dense model with parameters compared against an MoE model activating of its parameters per token. What this ratio hides is exactly what a compute-only comparison always hides: memory. Serving an MoE model requires holding all parameters resident — on one device or, more often, sharded across…
ense occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →oE occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →ense occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →ctive occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Inference Economics
This equation gives an approximation: it relates the quantities while allowing an approximation.
For inference cost specifically, what matters is that per-token compute tracks activated parameters, not total parameters. Approximating compute per token as proportional to the parameter count actually touched [ 13 , 11 ] , the ratio of dense to MoE compute at equal activated size is approximately . for a dense model with parameters compared against an MoE model activating of its parameters per token. What this ratio hides is exactly what a compute-only comparison always hides: memory. Serving an MoE model requires holding all parameters resident — on one device or, more often, sharded across…
Equation guide → · Article →