← All parts of this equation

Equation 5 · Part 7 · From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation

subscript

(IDF)k=⌈log⁡2n⌉−⌈log⁡2dk⌉+1,(\mathrm{IDF})_k = \lceil \log_2 n \rceil - \lceil \log_2 d_k \rceil + 1,
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

Before any of this involved learning, it involved geometry. Salton, Wong, and Yang proposed representing each document as a vector of weighted index terms and ranking documents by the similarity of their vectors to a query vector, arguing that a well-separated document space — one where unrelated documents sit far apart — should correspond to better retrieval performance than a densely packed one [ 1 ] . Their paper is worth reading in the original rather than through summary, because the term-weighting scheme it specifies is exactly the ancestor of what every later retriever, sparse or dense, still does: score a term by how often it occurs locally and how rare it is globally. They define…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.