← Mathematical compendium

Published equation contexts

Ptotal≈Pshared+E⋅PexpertP_{\mathrm{total}} \approx P_{\mathrm{shared}} + E \cdot P_{\mathrm{expert}}

Why this formula appears here

where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are PtotalP_{\mathrm{total}} and PactiveP_{\mathrm{active}} under a router…

Read the full article-specific guide →

Read the representative guide

EE

Symbol E

E is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ptotal≈Pshared+E⋅PexpertP_{\mathrm{total}} \approx P_{\mathrm{shared}} + E \cdot P_{\mathrm{expert}}

Equation 6 · Open Models

A History of Llama and the Open-Weight AI Movement

This equation gives an approximation: it relates the quantities while allowing an approximation.

where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are PtotalP_{\mathrm{total}} and PactiveP_{\mathrm{active}} under a router…

Meanings in this article

Equation guide → · Article →