Equation 2 · The Economics, Energy, and Physical Limits of Claude Code and Agentic Development Tools
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol t_token
oken is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol N_active
ctive occurs above the fraction bar. The numerator is divided by the entire denominator below it.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Denominator: BW
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
What the article says around this equation
Underneath all of the pricing mechanics sits a constraint pricing cannot remove: generating one token from a large language model is, for most of a request, bound by how fast bytes can move through accelerator memory rather than by how much arithmetic the accelerator can perform. Producing each token requires reading the model’s active weights and the accumulated key-value cache for that request from memory once; the arithmetic performed per byte read is low, so decoding is memory-bandwidth-bound while only the initial processing of a long prompt is compute-bound, a distinction Pope and colleagues formalized in their analysis of transformer inference efficiency [ 15 ] . Approximately,…
Read the full surrounding passage
Underneath all of the pricing mechanics sits a constraint pricing cannot remove: generating one token from a large language model is, for most of a request, bound by how fast bytes can move through accelerator memory rather than by how much arithmetic the accelerator can perform. Producing each token requires reading the model’s active weights and the accumulated key-value cache for that request from memory once; the arithmetic performed per byte read is low, so decoding is memory-bandwidth-bound while only the initial processing of a long prompt is compute-bound, a distinction Pope and colleagues formalized in their analysis of transformer inference efficiency [ 15 ] . Approximately, . where N-active is the number of parameters actively read per token, beta the bytes each occupies at the serving precision, KV of c the memory footprint of the key-value cache at context length c, and BW the accelerator’s memory bandwidth. Nothing about pricing changes this floor; it only changes which side of it a given task sits on. Kwon and colleagues showed that KV of c is frequently larger than it needs to be in practice, because conventional cache allocation reserves memory contiguously and wastes much of it to fragmentation, motivating the PagedAttention scheme that manages key-value memory the way an operating system manages virtual memory, in fixed-size pages rather than one contiguous block per request [ 16 ] . That is an engineering fix for waste inside the bound, not a repeal of the bound itself: a longer agentic session, carrying more context turn over turn in exactly the way the earlier accounting model describes, grows the key-value footprint and therefore grows the minimum time every subsequent token in that session can take, independent of price.
Sources cited in the surrounding passage
- [15] Efficiently Scaling Transformer Inference ↗
- [16] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Economics, Energy, and Physical Limits of Claude Code and Agentic Development Tools