← All parts of this equation

Equation 5 · Part 2 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

Symbol n

(IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1,(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1,
nn

What this part means

n is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

n is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…ancestor of what every later retriever, sparse or dense, still does: score a term by how often it occurs locally and how rare it is globally. They define the inverse document frequency of a term k , for a collection of n documents in which k appears in dkd_k of them, as (IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1. and combine it multiplicatively with raw term frequency so that a term scores highest when it occurs often in one document but rarely across the collection. Evaluated on three test collections in aerodynamics, medicine, and…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.