← Back to article

Equation 1 · How Benchmark Contamination Actually Works in Agentic Evaluation

What does this equation mean?

sobserved=scapability+Δleak+Δchecker+Δfingerprints_{\text{observed}} = s_{\text{capability}} + \Delta_{\text{leak}} + \Delta_{\text{checker}} + \Delta_{\text{fingerprint}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationss_capability + Delta_leak + Delta_checker + Delta_fingerprint
Result or conditions_observed
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

sobserveds_{\text{observed}}

Symbol s_observed

sos_observed is part of the quantity the equation computes from the expression on the right.

Understand this part →

scapabilitys_{\text{capability}}

Symbol s_capability

the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class.

Understand this part →

Δleak\Delta_{\text{leak}}

Symbol Delta_leak

inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer.

Understand this part →

Δchecker\Delta_{\text{checker}}

Symbol Delta_checker

inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on.

Understand this part →

Δfingerprint\Delta_{\text{fingerprint}}

Symbol Delta_fingerprint

a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: sobserved=scapability+Δleak+Δchecker+Δfingerprints_{\text{observed}} = s_{\text{capability}} + \Delta_{\text{leak}} + \Delta_{\text{checker}} + \Delta_{\text{fingerprint}}. scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather…
Read the full surrounding passage
It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: sobserved=scapability+Δleak+Δchecker+Δfingerprints_{\text{observed}} = s_{\text{capability}} + \Delta_{\text{leak}} + \Delta_{\text{checker}} + \Delta_{\text{fingerprint}}. scapabilitys_{\text{capability}} is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. Δleak\Delta_{\text{leak}} is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. Δchecker\Delta_{\text{checker}} is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on. Δfingerprint\Delta_{\text{fingerprint}} is a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task. A single-turn benchmark has to worry about something close to Δleak\Delta_{\text{leak}} alone. An agentic one has to worry about all three, and — this is the part that makes the topic worth its own article — the three interact, because a capable-enough policy can use the environment’s own tools to go looking for Δleak\Delta_{\text{leak}} on purpose.

Read the equation in its article →

For background, read the article’s source list.

Return to How Benchmark Contamination Actually Works in Agentic Evaluation

See this formula across 1 published context →

Browse the mathematical compendium →