← All parts of this equation

Equation 6 · Part 5 · A History of Llama and the Open-Weight AI Movement

≈

Ptotal≈Pshared+E⋅PexpertP_{\mathrm{total}} \approx P_{\mathrm{shared}} + E \cdot P_{\mathrm{expert}}
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are PtotalP_{\mathrm{total}} and PactiveP_{\mathrm{active}} under a router…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.