Symbol P_active
ctive is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
For a mixture-of-experts layer routing each token to one shared expert plus a fixed number of routed experts, the active parameter count per token is approximately . where is the always-active shared-expert capacity, is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly + E , while inference compute scales with , not . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400…
ctive is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Open Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
For a mixture-of-experts layer routing each token to one shared expert plus a fixed number of routed experts, the active parameter count per token is approximately . where is the always-active shared-expert capacity, is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly + E , while inference compute scales with , not . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400…