← All parts of this equation

Equation 8 · Part 4 · How Benchmark Contamination Actually Works in Agentic Evaluation

=

passk(t)=∏i=1k1 ⁣[trial i on task t succeeds],\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right],
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a passks^k metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption passks^k makes, and the empty-response exploit breaks exactly it: passk(t)=∏i=1k1 ⁣[trial i on task t succeeds]\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right]. averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.