← All parts of this equation

Equation 8 · Part 5 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

Symbol k_1

score(D,Q)=∑qi∈QIDF(qi)⋅f(qi,D)⋅(k1+1)f(qi,D)+k1⋅(1−b+b⋅∣D∣avgdl),\mathrm{score}(D, Q) = \sum_{q_i \in Q} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \dfrac{|D|}{\mathrm{avgdl}}\right)},
k1k_1

What this part means

k1k_1 is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

k1k_1 is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…a document D scores against a query Q as score(D,Q)=∑qi∈QIDF(qi)⋅f(qi,D)⋅(k1+1)f(qi,D)+k1⋅(1−b+b⋅∣D∣avgdl)\mathrm{score}(D, Q) = \sum_{q_i \in Q} \mathrm{IDF}(q_i) \cdot \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \dfrac{|D|}{\mathrm{avgdl}}\right)}. where f(qiq_i, D) is the frequency of query term qiq_i in D , |D| is the document’s length, avgdl\mathrm{avgdl} is the average document length in the collection, and k1k_1 and b are tuned constants controlling term-frequency saturation and length normalization respectively. The saturation term is the substantive advance over a raw term-frequency-times-IDF score: a term’s contribution grows quickly at first and then flattens, so a document that happens to repeat a…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.