Equation 16 · Measuring What a RAG System Retrieves, Not Just What It Answers
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →Numerator: claims supported by the retrieved context
The complete quantity above the fraction bar.
Denominator: total claims in the response
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Ragas is a widely used reference-free evaluation framework built explicitly around this separation: it scores retrieval effectiveness, the faithfulness with which a model uses retrieved passages, and generation quality, without requiring a human-written ground-truth answer for every query [ 4 ] . Its faithfulness metric is defined mechanically rather than as a vague notion of “sounding grounded”: a response is decomposed into individual factual claims, each claim is checked against the retrieved context to see whether it can be inferred from it, and the score is the resulting ratio, as the framework’s own documentation states [ 5 ] : . Notice what this formula does and…
Read the full surrounding passage
Ragas is a widely used reference-free evaluation framework built explicitly around this separation: it scores retrieval effectiveness, the faithfulness with which a model uses retrieved passages, and generation quality, without requiring a human-written ground-truth answer for every query [ 4 ] . Its faithfulness metric is defined mechanically rather than as a vague notion of “sounding grounded”: a response is decomposed into individual factual claims, each claim is checked against the retrieved context to see whether it can be inferred from it, and the score is the resulting ratio, as the framework’s own documentation states [ 5 ] : . Notice what this formula does and does not depend on. It does not depend on whether the retrieved context was the right context — a response can be perfectly faithful to a passage that answers a different question than the one asked, and the faithfulness score will not fall. It depends only on the relationship between the written output and whatever was placed in front of the model. Answer relevance and groundedness metrics, which the same documentation lists alongside faithfulness, complete the picture from the other side: relevance checks whether the response actually addresses the question rather than merely being consistent with the context, and groundedness-style checks look for unsupported additions the model introduced beyond what any retrieved passage stated [ 5 ] .
Sources cited in the surrounding passage
- [4] Ragas: Automated Evaluation of Retrieval Augmented Generation ↗
- [5] Faithfulness: RAGAS Metrics Documentation ↗
These citations give research context. Read each source to check which claims it supports.
Return to Measuring What a RAG System Retrieves, Not Just What It Answers