Equation 4 · A History of Llama and the Open-Weight AI Movement
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the number of routed experts activated per token. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where is the always-active shared-expert capacity, is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly + E , while inference compute scales with , not . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are and under a router…
Read the full surrounding passage
where is the always-active shared-expert capacity, is the parameter count of one routed expert, and k is the number of routed experts activated per token. Total capacity scales with the number of experts E available to the router, roughly + E , while inference compute scales with , not . This is the exact mechanism behind Meta’s headline figures: Maverick’s 400 billion total parameters and 17 billion active parameters are not two different measurements of the same quantity, they are and under a router with 128 available experts and a small k , and the entire commercial argument for the architecture — a much larger knowledge store served at close to small-model inference cost — depends on that gap holding up under real traffic rather than only in the launch announcement.
Sources cited in the article section
- [15] The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation ↗
- [16] Meta's Benchmarks for Its New AI Models Are a Bit Misleading ↗
- [17] Meta's 'Vanilla' Maverick AI Model Ranks Below Rivals on a Popular Chat Benchmark ↗
These citations give research context. Read each source to check which claims it supports.
Return to A History of Llama and the Open-Weight AI Movement