← Back to article

Equation 22 · Measuring What a RAG System Retrieves, Not Just What It Answers

What does this equation mean?

RR

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

RR

Symbol R

the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The reverse failure sits inside the first term. P(correct\text{correct} ∣\mid R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved…
Read the full surrounding passage
The reverse failure sits inside the first term. P(correct\text{correct} ∣\mid R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved documents, and at resisting counterfactual content placed in the context [ 8 ] . Negative rejection failure is the second term in the equation made concrete from the opposite direction: a system that answers confidently when it should abstain is a generation failure that a retrieval metric, scored only against relevance judgements, never sees.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Measuring What a RAG System Retrieves, Not Just What It Answers

Browse the mathematical compendium →