← Mathematical compendium

Published equation contexts

P(correct∣R)P(\text{correct} \mid R)

Why this formula appears here

The reverse failure sits inside the first term. P(correct\text{correct} ∣\mid R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved…

Read the full article-specific guide →

Read the representative guide

PP

Symbol P

P is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
RR

Symbol R

the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

P(correct∣R)P(\text{correct} \mid R)

Equation 21 · AI Agents & Systems

Measuring What a RAG System Retrieves, Not Just What It Answers

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The reverse failure sits inside the first term. P(correct\text{correct} ∣\mid R) can be low even when R holds — the right evidence was retrieved and still the generator wrote an unfaithful, evasive, or contradicted answer — and a retrieval-only evaluation, which never inspects generation at all, cannot detect this. Chen and colleagues benchmarked RAG systems along four separated axes — noise robustness, negative rejection, information integration, and counterfactual robustness — and found that while models tolerate some retrieved noise, they struggle badly at declining to answer when the retrieved context does not actually contain the answer, at integrating information across multiple retrieved…

Meanings in this article

  • RR: the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
Equation guide → · Article →