← All parts of this equation

Equation 8 · Part 3 · How Benchmark Contamination Actually Works in Agentic Evaluation

Symbol i

passk(t)=∏i=1k1 ⁣[trial i on task t succeeds],\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right],
ii

What this part means

i appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.

Its job in the formula

i appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a passks^k metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption passks^k makes, and the empty-response exploit breaks exactly it: passk(t)=∏i=1k1 ⁣[trial i on task t succeeds]\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right]. averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.