← Mathematical compendium

Published equation contexts

cfull(M)=Θ(M),cretrieve(M)=O(log⁡M)+Θ(k),k≪Mc_{\text{full}}(M) = \Theta(M), \qquad c_{\text{retrieve}}(M) = O(\log M) + \Theta(k), \qquad k \ll M

Why this formula appears here

The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k ≪\ll M tokens actually returned:…

Read the full article-specific guide →

Read the representative guide

cfullc_{\text{full}}

Symbol c_full

cfc_full is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
cretrievec_{\text{retrieve}}

Symbol c_retrieve

crc_retrieve is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

cfull(M)=Θ(M),cretrieve(M)=O(log⁡M)+Θ(k),k≪M.c_{\text{full}}(M) = \Theta(M), \qquad c_{\text{retrieve}}(M) = O(\log M) + \Theta(k), \qquad k \ll M.

Equation 6 · AI Agents & Systems

RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k ≪\ll M tokens actually returned:…

Meanings in this article

  • MM: the number of tokens.
Equation guide → · Article →