Equation 15 · Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hatp
hatp occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol n
n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
What the article says around this equation
4. Insufficient trial counts producing false confidence. A success rate reported from a small number of trials carries a much wider margin of error than the single reported number suggests, and the interval widens fast as the trial count shrinks. For a naive binomial estimate — treating each trial as an independent coin flip with unknown success probability p — the 95 percent confidence interval around an observed rate from n trials is approximately . At ten trials and an observed 90 percent success rate, that formula gives an interval of roughly plus or minus nineteen percentage points — an agent genuinely succeeding anywhere from about 71 percent to essentially…
Read the full surrounding passage
4. Insufficient trial counts producing false confidence. A success rate reported from a small number of trials carries a much wider margin of error than the single reported number suggests, and the interval widens fast as the trial count shrinks. For a naive binomial estimate — treating each trial as an independent coin flip with unknown success probability p — the 95 percent confidence interval around an observed rate from n trials is approximately . At ten trials and an observed 90 percent success rate, that formula gives an interval of roughly plus or minus nineteen percentage points — an agent genuinely succeeding anywhere from about 71 percent to essentially 100 percent of the time is statistically indistinguishable from the ten runs actually observed. Evan Miller, writing at Anthropic, makes this the central practical argument of a 2024 paper on evaluation statistics: he recommends that “new evals should contain at least 1,000 questions in order to have good signaling ability,” and works a concrete power calculation showing that detecting a genuine three-percentage-point difference between two systems, at conventional statistical standards (80 percent power, 5 percent significance), requires approximately 969 questions under reasonable assumptions [ 8 ] . Most agent benchmarks reported in the field — SWE-bench Verified’s 500 instances, τ-bench’s much smaller per-domain task sets — sit well below that figure, which does not make them worthless, but does mean that a one- or two-point difference between two systems on such a suite is frequently statistical noise dressed as a finding.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly