← Mathematical compendium

Published equation contexts

Ttok≥WBT_{\mathrm{tok}} \ge \frac{W}{B}

Why this formula appears here

Autoregressive decoding sits at the far left of that graph. Generating one token requires reading essentially every weight and the accumulated key-value cache once, and performing roughly two arithmetic operations per parameter. At one byte per weight the intensity is about two operations per byte; at four bits per weight, about four. Pope and colleagues formalised this partitioning problem for large transformers and showed how latency, throughput and cost trade off under different sharding strategies, with generation and prefill behaving as different regimes [ 3 ] . The bound that matters on a device follows immediately: if W bytes of weights and cache must cross the memory bus for each…

Read the full article-specific guide →

Read the representative guide

TtokT_{\mathrm{tok}}

Symbol T_tok

TtT_tok is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ttok≥WBT_{\mathrm{tok}} \ge \frac{W}{B}

Equation 8 · Small & Edge Models

The Deployment Envelope: Small Models Where the Power Is Not

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Autoregressive decoding sits at the far left of that graph. Generating one token requires reading essentially every weight and the accumulated key-value cache once, and performing roughly two arithmetic operations per parameter. At one byte per weight the intensity is about two operations per byte; at four bits per weight, about four. Pope and colleagues formalised this partitioning problem for large transformers and showed how latency, throughput and cost trade off under different sharding strategies, with generation and prefill behaving as different regimes [ 3 ] . The bound that matters on a device follows immediately: if W bytes of weights and cache must cross the memory bus for each…

Meanings in this article

  • WW: the halving the bits per weight halves.
  • BB: the achievable memory bandwidth and I the operational intensity of the kernel.
Equation guide → · Article →