← Back to article

Equation 19 · Measuring What a RAG System Retrieves, Not Just What It Answers

What does this equation mean?

P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

PP

Symbol P

P is part of the quantity the equation computes from the expression on the right.

Understand this part →

RR

Symbol R

the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}. An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P(correct\text{correct} ∣\mid ¬\lnot R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already…
Read the full surrounding passage
The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}. An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P(correct\text{correct} ∣\mid ¬\lnot R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was supposedly built around contributed nothing to the result.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Measuring What a RAG System Retrieves, Not Just What It Answers

See this formula across 1 published context →

Browse the mathematical compendium →