← Back to article

Equation 6 · RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

What does this equation mean?

cfull(M)=Θ(M),cretrieve(M)=O(log⁡M)+Θ(k),k≪M.c_{\text{full}}(M) = \Theta(M), \qquad c_{\text{retrieve}}(M) = O(\log M) + \Theta(k), \qquad k \ll M.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsTheta(M), qquad c_retrieve(M) = O(log M) + Theta(k), qquad k ll M
Result or conditionc_full(M)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cfullc_{\text{full}}

Symbol c_full

cfc_full is part of the quantity the equation computes from the expression on the right.

Understand this part →

MM

Symbol M

the number of tokens.

Understand this part →

Θ\Theta

Symbol Theta

Theta is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

cretrievec_{\text{retrieve}}

Symbol c_retrieve

crc_retrieve is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

OO

Symbol O

O is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

kk

Symbol k

k is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k ≪\ll M tokens actually returned:…
Read the full surrounding passage
The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k ≪\ll M tokens actually returned: cfull(M)=Θ(M),cretrieve(M)=O(log⁡M)+Θ(k),k≪Mc_{\text{full}}(M) = \Theta(M), \qquad c_{\text{retrieve}}(M) = O(\log M) + \Theta(k), \qquad k \ll M. As a corpus grows, the two costs diverge regardless of how cheap or how large any single model’s window becomes, because M keeps growing too. This is the economic argument for why explicit retrieval persists even in scenarios where verifiability is weak: corpus growth, not just model quality, keeps a bounded fetch cheaper than a full paste. It says nothing about whether that fetch is checkable, which is exactly why Axis B has to be argued separately rather than assumed to follow from Axis A.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

See this formula across 1 published context →

Browse the mathematical compendium →