Equation 19 · Measuring What a RAG System Retrieves, Not Just What It Answers
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol P
P is part of the quantity the equation computes from the expression on the right.
Symbol R
the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as . An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already…
Read the full surrounding passage
The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as . An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was supposedly built around contributed nothing to the result.
Sources cited in the article section
- [8] Benchmarking Large Language Models in Retrieval-Augmented Generation ↗
- [9] CRAG: Comprehensive RAG Benchmark ↗
These citations give research context. Read each source to check which claims it supports.
Return to Measuring What a RAG System Retrieves, Not Just What It Answers