← All parts of this equation

Equation 5 · Part 1 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

Symbol k

(IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1,(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1,
kk

What this part means

k is part of the quantity the equation computes from the expression on the right.

Its job in the formula

k is part of the quantity the equation computes from the expression on the right.

The passage around this formula

…specifies is exactly the ancestor of what every later retriever, sparse or dense, still does: score a term by how often it occurs locally and how rare it is globally. They define the inverse document frequency of a term k , for a collection of n documents in which k appears in dkd_k of them, as (IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1. and combine it multiplicatively with raw term frequency so that a term scores highest when it occurs often in one document but rarely across the collection. Evaluated on three test collections in…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.