← All parts of this equation

Equation 8 · Part 5 · How Benchmark Contamination Actually Works in Agentic Evaluation

subscript

passk(t)=∏i=1k1 ⁣[trial i on task t succeeds],\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right],
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a passks^k metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption passks^k makes, and the empty-response exploit breaks exactly it: passk(t)=∏i=1k1 ⁣[trial i on task t succeeds]\text{pass}^k(t) = \prod_{i=1}^{k} \mathbb{1}\!\left[\text{trial } i \text{ on task } t \text{ succeeds}\right]. averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.