Equation 1 · What Actually Happens Between a Query and an Answer in RAG
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol D
D is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol Q
Q is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol i
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol n
n appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol q_i
is one of the signed contributions combined to compute the quantity on the left.
Symbol f
f is one of the signed contributions combined to compute the quantity on the left.
Symbol k_1
is one of the signed contributions combined to compute the quantity on the left.
Symbol b
b occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Starting index or lower bound: i=1
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Ending index or upper bound: n
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Denominator: f(q_i, D) + k_1(1 - b + bdfrac|D|avgdl)
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Lexical retrieval scores a document by how well its terms overlap the query’s terms, weighted by how rare each term is and normalised for document length. The canonical scoring function, still the default first-stage ranker across a large share of production search two decades after its formulation, is BM25: . Robertson and Zaragoza’s account of the probabilistic relevance framework behind this formula is worth reading past the equation for one design choice it exposes: the term-frequency component saturates rather than growing linearly, so a document repeating a query term fifty times is scored only marginally higher than one repeating it five times, and the…
Read the full surrounding passage
Lexical retrieval scores a document by how well its terms overlap the query’s terms, weighted by how rare each term is and normalised for document length. The canonical scoring function, still the default first-stage ranker across a large share of production search two decades after its formulation, is BM25: . Robertson and Zaragoza’s account of the probabilistic relevance framework behind this formula is worth reading past the equation for one design choice it exposes: the term-frequency component saturates rather than growing linearly, so a document repeating a query term fifty times is scored only marginally higher than one repeating it five times, and the length-normalisation term b trades off penalising long documents against rewarding genuinely comprehensive ones [ 1 ] . BM25 has no notion of meaning. A query for “cannot authenticate” and a document about “login failure” share no terms and score zero, regardless of how obviously related a person would find them.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Actually Happens Between a Query and an Answer in RAG