← All parts of this equation

Equation 1 · Part 15 · What Actually Happens Between a Query and an Answer in RAG

superscript

score(D,Q)=∑i=1nIDF(qi)⋅f(qi,D) (k1+1)f(qi,D)+k1(1−b+b ∣D∣avgdl)\mathrm{score}(D,Q) = \sum_{i=1}^{n} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D)\,(k_1+1)}{f(q_i, D) + k_1\left(1 - b + b\,\dfrac{|D|}{\mathrm{avgdl}}\right)}
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Lexical retrieval scores a document by how well its terms overlap the query’s terms, weighted by how rare each term is and normalised for document length. The canonical scoring function, still the default first-stage ranker across a large share of production search two decades after its formulation, is BM25: score(D,Q)=∑i=1nIDF(qi)⋅f(qi,D) (k1+1)f(qi,D)+k1(1−b+b ∣D∣avgdl)\mathrm{score}(D,Q) = \sum_{i=1}^{n} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D)\,(k_1+1)}{f(q_i, D) + k_1\left(1 - b + b\,\dfrac{|D|}{\mathrm{avgdl}}\right)}. Robertson and Zaragoza’s account of the probabilistic relevance framework behind this formula is worth reading past the equation for one design choice it exposes: the term-frequency component saturates rather than growing linearly, so a document repeating a query term fifty times is scored only marginally higher than one repeating it five times, and the…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.