Equation 14 · Measuring What a RAG System Retrieves, Not Just What It Answers
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol k
the DCG of the ideal ordering, so the ratio is bounded near one regardless of how many relevant documents exist for a given query.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where is the graded relevance of the result at position i and k is the DCG of the ideal ordering, so the ratio is bounded near one regardless of how many relevant documents exist for a given query. The logarithmic discount encodes a specific judgement about attention: a relevant document at position one is worth far more than the same document at position ten, and nDCG is the standard way the information-retrieval literature makes that judgement quantitative. Heterogeneous retrieval benchmarks built to compare retrievers across many domains at once — BEIR evaluated ten lexical, sparse, dense, late-interaction and re-ranking systems across eighteen public datasets…
Read the full surrounding passage
where is the graded relevance of the result at position i and k is the DCG of the ideal ordering, so the ratio is bounded near one regardless of how many relevant documents exist for a given query. The logarithmic discount encodes a specific judgement about attention: a relevant document at position one is worth far more than the same document at position ten, and nDCG is the standard way the information-retrieval literature makes that judgement quantitative. Heterogeneous retrieval benchmarks built to compare retrievers across many domains at once — BEIR evaluated ten lexical, sparse, dense, late-interaction and re-ranking systems across eighteen public datasets and found BM25 a robust baseline that dense retrievers frequently underperformed out of domain — rely on exactly these rank-based metrics to make that comparison possible without ever generating an answer [ 2 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Measuring What a RAG System Retrieves, Not Just What It Answers