← All parts of this equation

Equation 7 · Part 1 · RAG in 2035: Four Scenarios, Their Signals, and What Would Falsify Them

Symbol M

MM
MM

What this part means

the number of tokens.

Its job in the formula

M is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

Pasting a whole corpus of size M into context puts a floor under n near M itself; retrieving costs roughly the price of an index lookup plus the price of reasoning over the k ≪\ll M tokens actually returned: cfull(M)c_{\text{full}}(M) = Θ(M)\Theta(M), \qquad cretrieve(M)c_{\text{retrieve}}(M) = O(log⁡\log M) + Θ(k)\Theta(k), \qquad k ≪\ll M.

The passage around this formula

As a corpus grows, the two costs diverge regardless of how cheap or how large any single model’s window becomes, because M keeps growing too. This is the economic argument for why explicit retrieval persists even in scenarios where verifiability is weak: corpus growth, not just model quality, keeps a bounded fetch cheaper than a full paste. It says nothing about whether that fetch is checkable, which is exactly why Axis B has to be argued separately rather than assumed to follow from Axis A.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.