Equation 22 · Measuring What a RAG System Retrieves, Not Just What It Answers
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol R
the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The reverse failure sits inside the first term. P( R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved…
Read the full surrounding passage
The reverse failure sits inside the first term. P( R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved documents, and at resisting counterfactual content placed in the context [ 8 ] . Negative rejection failure is the second term in the equation made concrete from the opposite direction: a system that answers confidently when it should abstain is a generation failure that a retrieval metric, scored only against relevance judgements, never sees.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Measuring What a RAG System Retrieves, Not Just What It Answers