Symbol P
P is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was…
P is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →the event that the retriever’s top- k results contain sufficient evidence to answer the query, and let “correct” mean the generated answer is judged faithful and relevant.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 20 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
An end-to-end accuracy score measures the left-hand side only. The two terms on the right are exactly what generation metrics and retrieval metrics separately illuminate, and the second term is the one an aggregate score cannot see at all. P( R) is the probability the model answers correctly despite the retriever having failed — because the underlying language model already carried the answer parametrically, because the question was easy enough to guess, or, in the worst case, because the benchmark’s own answer had leaked into the model’s training data. Whenever that term is large, an end-to-end score can look excellent while the retrieval component the system was…
Equation 23 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
The mechanism is not exotic. A benchmark released as a public dataset, a paper, or a web page is exactly the kind of content large-scale pretraining crawls are built to ingest. A survey of benchmark data contamination in large language models lays out the resulting problem directly: leading models trained on web-scale corpora can inadvertently incorporate benchmark data into their training sets, which inflates measured performance in ways that do not reflect genuine capability, and the survey catalogues detection methods and alternative assessment strategies developed specifically to counter it [ 12 ] . For RAG evaluation the consequence is sharper than for a plain language-model benchmark,…