← Back to article

Equation 8 · Model Systems in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

What does this equation mean?

ttoken≳W+KB,t_{\mathrm{token}} \gtrsim \frac{W + K}{B},

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ttokent_{\mathrm{token}}

Symbol t_token

ttt_token is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

WW

Symbol W

the weight bytes touched and K the cache bytes.

Understand this part →

KK

Symbol K

the cache bytes.

Understand this part →

BB

Symbol B

B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

W+KW + K

Numerator: W + K

The complete quantity above the fraction bar.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

What the article says around this equation

with gFg_F and gBg_B the annual growth factors. When gFg_F > gBg_B , I∗I^{*} grows without bound, and the batch size required to keep the arithmetic units busy grows with it. Autoregressive decoding sits on the wrong side of this: generating one token requires streaming the weights and the accumulated key–value cache, so decode time is bounded below by ttoken≳W+KBt_{\mathrm{token}} \gtrsim \frac{W + K}{B}. where W is the weight bytes touched and K the cache bytes. Pope and colleagues formalised the partitioning analysis behind this and showed how latency, throughput and cost trade against one another under different sharding layouts [ 12 ] . The mitigations are well documented and all attack the numerator or the batch: grouped-query…
Read the full surrounding passage
with gFg_F and gBg_B the annual growth factors. When gFg_F > gBg_B , I∗I^{*} grows without bound, and the batch size required to keep the arithmetic units busy grows with it. Autoregressive decoding sits on the wrong side of this: generating one token requires streaming the weights and the accumulated key–value cache, so decode time is bounded below by ttoken≳W+KBt_{\mathrm{token}} \gtrsim \frac{W + K}{B}. where W is the weight bytes touched and K the cache bytes. Pope and colleagues formalised the partitioning analysis behind this and showed how latency, throughput and cost trade against one another under different sharding layouts [ 12 ] . The mitigations are well documented and all attack the numerator or the batch: grouped-query attention shrinks the key–value cache by sharing key and value heads across query groups [ 14 ] ; PagedAttention removes fragmentation and over-reservation in cache allocation, reported at 2–4× throughput at equal latency [ 13 ] ; speculative decoding lets a cheap draft model propose tokens that the target verifies in parallel, provably without changing the sampled distribution [ 15 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Model Systems in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

See this formula across 1 published context →

Browse the mathematical compendium →