Equation 15 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
tuned constants controlling term-frequency saturation and length normalization respectively. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol b
tuned constants controlling term-frequency saturation and length normalization respectively.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where f(, D) is the frequency of query term in D , |D| is the document’s length, is the average document length in the collection, and and b are tuned constants controlling term-frequency saturation and length normalization respectively. The saturation term is the substantive advance over a raw term-frequency-times-IDF score: a term’s contribution grows quickly at first and then flattens, so a document that happens to repeat a query word fifty times does not dominate one that uses it three times in the right place. Decades later, this same function remains the default first-stage ranking method built into widely used open-source search engines, which is a strong…
Read the full surrounding passage
where f(, D) is the frequency of query term in D , |D| is the document’s length, is the average document length in the collection, and and b are tuned constants controlling term-frequency saturation and length normalization respectively. The saturation term is the substantive advance over a raw term-frequency-times-IDF score: a term’s contribution grows quickly at first and then flattens, so a document that happens to repeat a query word fifty times does not dominate one that uses it three times in the right place. Decades later, this same function remains the default first-stage ranking method built into widely used open-source search engines, which is a strong claim to make about any piece of 1990s software and is offered here as an observation about longevity rather than as evidence that nothing since has improved on it.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation