Equation 5 · How Benchmark Contamination Actually Works in Agentic Evaluation
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol Delta_fingerprint
a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior…
Read the full surrounding passage
is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. A single-turn benchmark has to worry about something close to alone. An agentic one has to worry about all three, and — this is the part that makes the topic worth its own article — the three interact, because a capable-enough policy can use the environment’s own tools to go looking for on purpose.
For background, read the article’s source list.
Return to How Benchmark Contamination Actually Works in Agentic Evaluation