← Back to article

Equation 24 · What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach

What does this equation mean?

tstep≥Bweights+Bkvβ,t_{\mathrm{step}} \ge \frac{B_{\mathrm{weights}} + B_{\mathrm{kv}}}{\beta},

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

tstept_{\mathrm{step}}

Symbol t_step

tst_step is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

BweightsB_{\mathrm{weights}}

Symbol B_weights

BwB_weights occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

BkvB_{\mathrm{kv}}

Symbol B_kv

BkB_kv occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

β\beta

Symbol β

β occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

Bweights+BkvB_{\mathrm{weights}} + B_{\mathrm{kv}}

Numerator: B_weights + B_kv

The complete quantity above the fraction bar.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

What the article says around this equation

Bandwidth-limited throughput is the regime where achieved FLOPS is irrelevant because β\beta ⋅\cdot I is the binding term. Autoregressive decoding in a served language model is the clearest case: generating a single token requires streaming the model weights and the accumulated key-value cache out of memory, and performs only a small number of operations per byte read. The time per decoding step obeys tstep≥Bweights+Bkvβt_{\mathrm{step}} \ge \frac{B_{\mathrm{weights}} + B_{\mathrm{kv}}}{\beta}. a floor set entirely by memory traffic, in which peak arithmetic does not appear. This is why batching improves throughput so dramatically — the same weight bytes are amortised across many sequences, raising I — and why it does not improve single-stream latency at all.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach

See this formula across 1 published context →

Browse the mathematical compendium →