Equation 1 · How Benchmark Contamination Actually Works in Agentic Evaluation
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol s_observed
bserved is part of the quantity the equation computes from the expression on the right.
Symbol s_capability
the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class.
Symbol Delta_leak
inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer.
Symbol Delta_checker
inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on.
Symbol Delta_fingerprint
a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: . is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather…
Read the full surrounding passage
It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: . is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. A single-turn benchmark has to worry about something close to alone. An agentic one has to worry about all three, and — this is the part that makes the topic worth its own article — the three interact, because a capable-enough policy can use the environment’s own tools to go looking for on purpose.
For background, read the article’s source list.
Return to How Benchmark Contamination Actually Works in Agentic Evaluation