Equation 8 · Part 7 · How Benchmark Contamination Actually Works in Agentic Evaluation
Starting index or lower bound: i=1
What this part means
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Its job in the formula
i=1 appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Starting index or lower bound: i=1→Article meaning
The passage around this formula
tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a pas metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption pas makes, and the empty-response exploit breaks exactly it: . averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…
Learn the underlying idea
Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.
Open the illustrated sums and products: repeat an operation over an index guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.