← Back to article

Equation 5 · How Benchmark Contamination Actually Works in Agentic Evaluation

What does this equation mean?

Δfingerprint\Delta_{\text{fingerprint}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Δfingerprint\Delta_{\text{fingerprint}}

Symbol Delta_fingerprint

a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. Δfingerprint\Delta_{\text{fingerprint}} is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior…
Read the full surrounding passage
scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. Δfingerprint\Delta_{\text{fingerprint}} is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. A single-turn benchmark has to worry about something close to Δleak\Delta_{\text{leak}} alone. An agentic one has to worry about all three, and — this is the part that makes the topic worth its own article — the three interact, because a capable-enough policy can use the environment’s own tools to go looking for Δleak\Delta_{\text{leak}} on purpose.

Read the equation in its article →

For background, read the article’s source list.

Return to How Benchmark Contamination Actually Works in Agentic Evaluation

Browse the mathematical compendium →