← All parts of this equation

Equation 5 · Part 3 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

Symbol d_k

(IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1,(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1,
dkd_k

What this part means

dkd_k is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

dkd_k is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…sparse or dense, still does: score a term by how often it occurs locally and how rare it is globally. They define the inverse document frequency of a term k , for a collection of n documents in which k appears in dkd_k of them, as (IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1. and combine it multiplicatively with raw term frequency so that a term scores highest when it occurs often in one document but rarely across the collection. Evaluated on three test collections in aerodynamics, medicine, and world affairs, replacing raw…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.