← All parts of this equation

Equation 24 · Part 6 · What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach

≥

tstep≥Bweights+Bkvβ,t_{\mathrm{step}} \ge \frac{B_{\mathrm{weights}} + B_{\mathrm{kv}}}{\beta},
≥

What this part means

Greater than or equal to.

Its job in the formula

Greater than or equal to.

The passage around this formula

Bandwidth-limited throughput is the regime where achieved FLOPS is irrelevant because β\beta ⋅\cdot I is the binding term. Autoregressive decoding in a served language model is the clearest case: generating a single token requires streaming the model weights and the accumulated key-value cache out of memory, and performs only a small number of operations per byte read. The time per decoding step obeys tstep≥Bweights+Bkvβt_{\mathrm{step}} \ge \frac{B_{\mathrm{weights}} + B_{\mathrm{kv}}}{\beta}. a floor set entirely by memory traffic, in which peak arithmetic does not appear. This is why batching improves throughput so dramatically — the same weight bytes are amortised across many sequences, raising I — and why it does not improve single-stream latency at all.

Read this part in the article →

Learn the underlying idea

An inequality compares values without claiming they are equal. It describes a range, threshold, or bound that a quantity may satisfy.

Open the illustrated inequalities: bounds and allowed ranges guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.