← Back to article

Equation 7 · How Benchmark Contamination Actually Works in Agentic Evaluation

What does this equation mean?

Δleak\Delta_{\text{leak}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Δleak\Delta_{\text{leak}}

Symbol Delta_leak

inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. Δfingerprint\Delta_{\text{fingerprint}} is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior…
Read the full surrounding passage
scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. Δfingerprint\Delta_{\text{fingerprint}} is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. A single-turn benchmark has to worry about something close to Δleak\Delta_{\text{leak}} alone. An agentic one has to worry about all three, and — this is the part that makes the topic worth its own article — the three interact, because a capable-enough policy can use the environment’s own tools to go looking for Δleak\Delta_{\text{leak}} on purpose.

Read the equation in its article →

For background, read the article’s source list.

Return to How Benchmark Contamination Actually Works in Agentic Evaluation

Browse the mathematical compendium →