← All parts of this equation

Equation 1 · Part 16 · What Actually Happens Between a Query and an Answer in RAG

Starting index or lower bound: i=1

score(D,Q)=∑i=1nIDF(qi)⋅f(qi,D) (k1+1)f(qi,D)+k1(1−b+b ∣D∣avgdl)\mathrm{score}(D,Q) = \sum_{i=1}^{n} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D)\,(k_1+1)}{f(q_i, D) + k_1\left(1 - b + b\,\dfrac{|D|}{\mathrm{avgdl}}\right)}
i=1i=1

What this part means

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Its job in the formula

i=1 appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

Lexical retrieval scores a document by how well its terms overlap the query’s terms, weighted by how rare each term is and normalised for document length. The canonical scoring function, still the default first-stage ranker across a large share of production search two decades after its formulation, is BM25: score(D,Q)=∑i=1nIDF(qi)⋅f(qi,D) (k1+1)f(qi,D)+k1(1−b+b ∣D∣avgdl)\mathrm{score}(D,Q) = \sum_{i=1}^{n} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D)\,(k_1+1)}{f(q_i, D) + k_1\left(1 - b + b\,\dfrac{|D|}{\mathrm{avgdl}}\right)}. Robertson and Zaragoza’s account of the probabilistic relevance framework behind this formula is worth reading past the equation for one design choice it exposes: the term-frequency component saturates rather than growing linearly, so a document repeating a query term fifty times is scored only marginally higher than one repeating it five times, and the…

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.