← Mathematical compendium

Published equation contexts

passdeterministick=pass1for every k\text{pass}^k_{\text{deterministic}} = \text{pass}^1 \quad \text{for every } k

Why this formula appears here

averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but inconsistent across resampled trials — a stochastic policy with true per-trial success probability p has passk\text{pass}^k ≈\approx pkp^k , which falls quickly as k grows. But a policy whose output on a given task is deterministic — an empty response, always, regardless of sampling — produces the identical transcript on every trial, so passdeterministick=pass1for every k\text{pass}^k_{\text{deterministic}} = \text{pass}^1 \quad \text{for every } k . A degenerate exploit is not merely invisible to pass@1; it is more invisible to passks^k, precisely because passks^k was built to reward consistency, and a fixed, checker-satisfying non-answer is the most consistent…

Read the full article-specific guide →

Read the representative guide

kdeterministick_{\text{deterministic}}

Symbol k_deterministic

kdk_deterministic is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

passdeterministick=pass1for every k.\text{pass}^k_{\text{deterministic}} = \text{pass}^1 \quad \text{for every } k .

Equation 12 · Model Evaluation

How Benchmark Contamination Actually Works in Agentic Evaluation

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but inconsistent across resampled trials — a stochastic policy with true per-trial success probability p has passk\text{pass}^k ≈\approx pkp^k , which falls quickly as k grows. But a policy whose output on a given task is deterministic — an empty response, always, regardless of sampling — produces the identical transcript on every trial, so passdeterministick=pass1for every k\text{pass}^k_{\text{deterministic}} = \text{pass}^1 \quad \text{for every } k . A degenerate exploit is not merely invisible to pass@1; it is more invisible to passks^k, precisely because passks^k was built to reward consistency, and a fixed, checker-satisfying non-answer is the most consistent…

Equation guide → · Article →