Equation 8 · Ten Failure Modes That Define Production RAG
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol k
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Some queries have no single supporting span, by construction. The answer requires combining a fact from one document with a fact from another, sometimes conditionally, and no amount of top- k tuning against a single-hop retriever fixes an architecture that was never asked to do the second half of that job. Tang and Yang built MultiHop-RAG specifically to test this, pairing a knowledge base with multi-hop queries whose answers require retrieving and reasoning across multiple supporting documents rather than one, and found that existing retrieval-augmented pipelines perform unsatisfactorily on it — both the embedding models used for retrieval and the language models used for reasoning over…
Read the full surrounding passage
Some queries have no single supporting span, by construction. The answer requires combining a fact from one document with a fact from another, sometimes conditionally, and no amount of top- k tuning against a single-hop retriever fixes an architecture that was never asked to do the second half of that job. Tang and Yang built MultiHop-RAG specifically to test this, pairing a knowledge base with multi-hop queries whose answers require retrieving and reasoning across multiple supporting documents rather than one, and found that existing retrieval-augmented pipelines perform unsatisfactorily on it — both the embedding models used for retrieval and the language models used for reasoning over what was retrieved struggled with the multi-hop case specifically, across every system-of-the-day tested including GPT-4-class models [ 9 ] . The failure has two independent causes stacked on top of each other: single-hop retrieval optimises for finding one relevant document per query, so it is not even trying to assemble the right set for a compositional question, and even when the full set is somehow present in context, chaining facts across documents that were written independently, in different styles, at different times, is a harder reasoning task than the single-document case the rest of this article implicitly assumes.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.