Equation 8 · Part 2 · How Benchmark Contamination Actually Works in Agentic Evaluation
Symbol t
What this part means
t is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Its job in the formula
t is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Full expression→Symbol t→Article meaning
The passage around this formula
tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a pas metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption pas makes, and the empty-response exploit breaks exactly it: . averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.