Equation 6 · RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol c_full
ull is part of the quantity the equation computes from the expression on the right.
Symbol Theta
Theta is one of the signed contributions combined to compute the quantity on the left.
Symbol c_retrieve
etrieve is one of the signed contributions combined to compute the quantity on the left.
Symbol O
O is one of the signed contributions combined to compute the quantity on the left.
Symbol k
k is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k M tokens actually returned:…
Read the full surrounding passage
The two axes are related but not reducible to one another, and the coupling has an economic root worth making explicit. Autoregressive decoding is bounded by how much has to be read from memory to produce each token — the weights, plus whatever context has accumulated — and Pope and colleagues formalized exactly this partitioning of cost between arithmetic and memory movement in transformer serving [ 5 ] . Write n for the number of tokens a system must hold in an answer’s working set. Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k M tokens actually returned: . As a corpus grows, the two costs diverge regardless of how cheap or how large any single model’s window becomes, because M keeps growing too. This is the economic argument for why explicit retrieval persists even in scenarios where verifiability is weak: corpus growth, not just model quality, keeps a bounded fetch cheaper than a full paste. It says nothing about whether that fetch is checkable, which is exactly why Axis B has to be argued separately rather than assumed to follow from Axis A.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them