← All parts of this equation

Equation 1 · Part 4 · A History of Llama and the Open-Weight AI Movement

Symbol P_expert

Pactive≈Pshared+k⋅Pexpert,P_{\mathrm{active}} \approx P_{\mathrm{shared}} + k \cdot P_{\mathrm{expert}},
PexpertP_{\mathrm{expert}}

What this part means

the parameter count of one routed expert.

Its job in the formula

PeP_expert is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token.

The passage around this formula

…each token to one shared expert plus a fixed number of routed experts, the active parameter count per token is approximately Pactive≈Pshared+k⋅PexpertP_{\mathrm{active}} \approx P_{\mathrm{shared}} + k \cdot P_{\mathrm{expert}}. where PsharedP_{\mathrm{shared}} is the always-active shared-expert capacity, PexpertP_{\mathrm{expert}} is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly PtotalP_{\mathrm{total}} ≈\approx PsharedP_{\mathrm{shared}} + E ⋅\cdot PexpertP_{\mathrm{expert}} , while…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.