← Back to article

Equation 8 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

What does this equation mean?

score(D,Q)=∑qi∈QIDF(qi)⋅f(qi,D)⋅(k1+1)f(qi,D)+k1⋅(1−b+b⋅∣D∣avgdl),\mathrm{score}(D, Q) = \sum_{q_i \in Q} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \dfrac{|D|}{\mathrm{avgdl}}\right)},

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withf(q_i, D) × (k_1 + 1)
Divide byf(q_i, D) + k_1 × (1 - b + b × dfrac|D|avgdl)
This relates toscore(D, Q)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

DD

Symbol D

D is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

QQ

Symbol Q

Q appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

qiq_i

Symbol q_i

qiq_i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

ff

Symbol f

the frequency of query term qiq_i in D , |D| is the document’s length, avgdl\mathrm{avgdl} is the average document length in the collection, and k1k_1.

Understand this part →

k1k_1

Symbol k_1

k1k_1 is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

bb

Symbol b

tuned constants controlling term-frequency saturation and length normalization respectively.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

qi∈Qq_i \in Q

Starting index or lower bound: q_i in Q

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

f(qi,D)⋅(k1+1)f(q_i, D) \cdot (k_1 + 1)

Numerator: f(q_i, D) × (k_1 + 1)

The complete quantity above the fraction bar.

Understand this part →

f(qi,D)+k1⋅(1−b+b⋅∣D∣avgdl)f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \dfrac{|D|}{\mathrm{avgdl}}\right)

Denominator: f(q_i, D) + k_1 × (1 - b + b × dfrac|D|avgdl)

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The Text REtrieval Conference, run annually by the U.S. National Institute of Standards and Technology from 1992 onward, gave information retrieval something the field had mostly lacked: a shared, blind evaluation on a common document set, repeated every year with published results. Robertson, Walker, and colleagues at City University London entered TREC-3 in 1994 with the Okapi system, reporting a term-weighting approach within a probabilistic relevance framework and applying it, among other extensions, to phrase weighting and to query expansion using terms drawn from an initial pilot search [ 2 ] . Over that and the following TREC rounds, the Okapi team’s tuning of term-frequency…
Read the full surrounding passage
The Text REtrieval Conference, run annually by the U.S. National Institute of Standards and Technology from 1992 onward, gave information retrieval something the field had mostly lacked: a shared, blind evaluation on a common document set, repeated every year with published results. Robertson, Walker, and colleagues at City University London entered TREC-3 in 1994 with the Okapi system, reporting a term-weighting approach within a probabilistic relevance framework and applying it, among other extensions, to phrase weighting and to query expansion using terms drawn from an initial pilot search [ 2 ] . Over that and the following TREC rounds, the Okapi team’s tuning of term-frequency saturation and document-length normalization converged on the specific scoring function that the field came to call BM25 — “Best Match 25,” after its position in a numbered sequence of variants the group had tried. Written in its now-standard form, a document D scores against a query Q as score(D,Q)=∑qi∈QIDF(qi)⋅f(qi,D)⋅(k1+1)f(qi,D)+k1⋅(1−b+b⋅∣D∣avgdl)\mathrm{score}(D, Q) = \sum_{q_i \in Q} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \dfrac{|D|}{\mathrm{avgdl}}\right)}. where f(qiq_i, D) is the frequency of query term qiq_i in D , |D| is the document’s length, avgdl\mathrm{avgdl} is the average document length in the collection, and k1k_1 and b are tuned constants controlling term-frequency saturation and length normalization respectively. The saturation term is the substantive advance over a raw term-frequency-times-IDF score: a term’s contribution grows quickly at first and then flattens, so a document that happens to repeat a query word fifty times does not dominate one that uses it three times in the right place. Decades later, this same function remains the default first-stage ranking method built into widely used open-source search engines, which is a strong claim to make about any piece of 1990s software and is offered here as an observation about longevity rather than as evidence that nothing since has improved on it.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

See this formula across 1 published context →

Browse the mathematical compendium →