Equation 20 · Measuring What a RAG System Retrieves, Not Just What It Answers
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol P
P is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol R
the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was…
Read the full surrounding passage
An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was supposedly built around contributed nothing to the result.
Sources cited in the article section
- [8] Benchmarking Large Language Models in Retrieval-Augmented Generation ↗
- [9] CRAG: Comprehensive RAG Benchmark ↗
These citations give research context. Read each source to check which claims it supports.
Return to Measuring What a RAG System Retrieves, Not Just What It Answers