← Mathematical compendium

Published equation contexts

P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}

Why this formula appears here

The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}. An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P(correct\text{correct} ∣\mid ¬\lnot R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already…

Read the full article-specific guide →

Read the representative guide

RR

Symbol R

the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}

Equation 19 · AI Agents & Systems

Measuring What a RAG System Retrieves, Not Just What It Answers

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The measurement structure explains the divergence directly. Let R be the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Total correctness decomposes as P(correct)  =  P(R) P(correct∣R)⏟grounded in retrieved evidence  +  P(¬R) P(correct∣¬R)⏟correct without sufficient retrieved evidenceP(\text{correct}) \;=\; \underbrace{P(R)\,P(\text{correct} \mid R)}_{\text{grounded in retrieved evidence}} \;+\; \underbrace{P(\lnot R)\,P(\text{correct} \mid \lnot R)}_{\text{correct without sufficient retrieved evidence}}. An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P(correct\text{correct} ∣\mid ¬\lnot R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already…

Meanings in this article

  • RR: the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
Equation guide → · Article →