← Mathematical compendium

Published equation contexts

ttoken≳W+KBt_{\mathrm{token}} \gtrsim \frac{W + K}{B}

Why this formula appears here

with gFg_F and gBg_B the annual growth factors. When gFg_F > gBg_B , I∗I^{*} grows without bound, and the batch size required to keep the arithmetic units busy grows with it. Autoregressive decoding sits on the wrong side of this: generating one token requires streaming the weights and the accumulated key–value cache, so decode time is bounded below by ttoken≳W+KBt_{\mathrm{token}} \gtrsim \frac{W + K}{B}. where W is the weight bytes touched and K the cache bytes. Pope and colleagues formalised the partitioning analysis behind this and showed how latency, throughput and cost trade against one another under different sharding layouts [ 12 ] . The mitigations are well documented and all attack the numerator or the batch: grouped-query…

Read the full article-specific guide →

Read the representative guide

ttokent_{\mathrm{token}}

Symbol t_token

ttt_token is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
BB

Symbol B

B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

ttoken≳W+KB,t_{\mathrm{token}} \gtrsim \frac{W + K}{B},

Equation 8 · Foundation Models

Model Systems in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

with gFg_F and gBg_B the annual growth factors. When gFg_F > gBg_B , I∗I^{*} grows without bound, and the batch size required to keep the arithmetic units busy grows with it. Autoregressive decoding sits on the wrong side of this: generating one token requires streaming the weights and the accumulated key–value cache, so decode time is bounded below by ttoken≳W+KBt_{\mathrm{token}} \gtrsim \frac{W + K}{B}. where W is the weight bytes touched and K the cache bytes. Pope and colleagues formalised the partitioning analysis behind this and showed how latency, throughput and cost trade against one another under different sharding layouts [ 12 ] . The mitigations are well documented and all attack the numerator or the batch: grouped-query…

Meanings in this article

  • WW: the weight bytes touched and K the cache bytes.
  • KK: the cache bytes.
Equation guide → · Article →