← Back to article

Equation 6 · A History of Llama and the Open-Weight AI Movement

What does this equation mean?

Ptotal≈Pshared+E⋅PexpertP_{\mathrm{total}} \approx P_{\mathrm{shared}} + E \cdot P_{\mathrm{expert}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

PtotalP_{\mathrm{total}}

Symbol P_total

the not.

Understand this part →

PsharedP_{\mathrm{shared}}

Symbol P_shared

the always-active shared-expert capacity.

Understand this part →

EE

Symbol E

E is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

PexpertP_{\mathrm{expert}}

Symbol P_expert

the parameter count of one routed expert.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are PtotalP_{\mathrm{total}} and PactiveP_{\mathrm{active}} under a router…
Read the full surrounding passage
where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are PtotalP_{\mathrm{total}} and PactiveP_{\mathrm{active}} under a router with 128 available experts and a small k , and the entire commercial argument for the architecture — a much larger knowledge store served at close to small-model inference cost — depends on that gap holding up under real traffic rather than only in the launch announcement.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A History of Llama and the Open-Weight AI Movement

See this formula across 1 published context →

Browse the mathematical compendium →