← All parts of this equation

Equation 8 · Part 1 · A History of Llama and the Open-Weight AI Movement

Symbol P_total

PtotalP_{\mathrm{total}}
PtotalP_{\mathrm{total}}

What this part means

the not.

Its job in the formula

PtP_total is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} .

The passage around this formula

…capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while inference compute scales with PactiveP_{\mathrm{active}} , not PtotalP_{\mathrm{total}} . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.