Equation 8 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol D
D is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol Q
Q appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol q_i
appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol f
the frequency of query term in D , |D| is the document’s length, is the average document length in the collection, and .
Symbol k_1
is one of the signed contributions combined to compute the quantity on the left.
Symbol b
tuned constants controlling term-frequency saturation and length normalization respectively.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Starting index or lower bound: q_i in Q
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Denominator: f(q_i, D) + k_1 × (1 - b + b × dfrac|D|avgdl)
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The Text REtrieval Conference, run annually by the U.S. National Institute of Standards and Technology from 1992 onward, gave information retrieval something the field had mostly lacked: a shared, blind evaluation on a common document set, repeated every year with published results. Robertson, Walker, and colleagues at City University London entered TREC-3 in 1994 with the Okapi system, reporting a term-weighting approach within a probabilistic relevance framework and applying it, among other extensions, to phrase weighting and to query expansion using terms drawn from an initial pilot search [ 2 ] . Over that and the following TREC rounds, the Okapi team’s tuning of term-frequency…
Read the full surrounding passage
The Text REtrieval Conference, run annually by the U.S. National Institute of Standards and Technology from 1992 onward, gave information retrieval something the field had mostly lacked: a shared, blind evaluation on a common document set, repeated every year with published results. Robertson, Walker, and colleagues at City University London entered TREC-3 in 1994 with the Okapi system, reporting a term-weighting approach within a probabilistic relevance framework and applying it, among other extensions, to phrase weighting and to query expansion using terms drawn from an initial pilot search [ 2 ] . Over that and the following TREC rounds, the Okapi team’s tuning of term-frequency saturation and document-length normalization converged on the specific scoring function that the field came to call BM25 — “Best Match 25,” after its position in a numbered sequence of variants the group had tried. Written in its now-standard form, a document D scores against a query Q as . where f(, D) is the frequency of query term in D , |D| is the document’s length, is the average document length in the collection, and and b are tuned constants controlling term-frequency saturation and length normalization respectively. The saturation term is the substantive advance over a raw term-frequency-times-IDF score: a term’s contribution grows quickly at first and then flattens, so a document that happens to repeat a query word fifty times does not dominate one that uses it three times in the right place. Decades later, this same function remains the default first-stage ranking method built into widely used open-source search engines, which is a strong claim to make about any piece of 1990s software and is offered here as an observation about longevity rather than as evidence that nothing since has improved on it.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation