← All parts of this equation

Equation 2 · Part 3 · The Economics, Energy, and Physical Limits of Claude Code and Agentic Development Tools

Symbol β

ttoken  ≳  2 Nactive β+KV(c)BW,t_{\mathrm{token}} \;\gtrsim\; \frac{2\,N_{\mathrm{active}}\,\beta + \mathrm{KV}(c)}{\mathrm{BW}},
β\beta

What this part means

the bytes each occupies at the serving precision.

Its job in the formula

β occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Where the article explains it

where N-active is the number of parameters actively read per token, beta the bytes each occupies at the serving precision, KV of c the memory footprint of the key-value cache at context length c, and BW the accelerator’s memory bandwidth.

The passage around this formula

Underneath all of the pricing mechanics sits a constraint pricing cannot remove: generating one token from a large language model is, for most of a request, bound by how fast bytes can move through accelerator memory rather than by how much arithmetic the accelerator can perform. Producing each token requires reading the model’s active weights and the accumulated key-value cache for that request from memory once; the arithmetic performed per byte read is low, so decoding is memory-bandwidth-bound while only the initial processing of a long prompt is compute-bound, a distinction Pope and colleagues formalized in their analysis of transformer inference efficiency [ 15 ] . Approximately,…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.