Equation 22 · Measuring AI Agent Reliability: What the Evidence Actually Supports
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol n
n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol z_α/2
z_α/2 occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol z_β
z_β occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol delta^2
delt occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Numerator: (z_α/2 + z_β)^2
The complete quantity above the fraction bar.
See an illustrated explanation →Denominator: 9delta^2
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
What the article says around this equation
Even when a team does report a rate rather than a single anecdote, a rate on its own understates how uncertain that number is. Evan Miller’s statistical treatment of language-model evaluation makes an argument borrowed from experimental science generally: an evaluation score is an estimate drawn from a finite, noisy sample, not a fact about the system, and it should be reported the way any other science reports a measurement — with an uncertainty attached [ 4 ] . His worked illustration is a useful gut check on how much data “enough” actually requires. To reliably detect an absolute difference of three percentage points between two systems, with conventional statistical standards for false…
Read the full surrounding passage
Even when a team does report a rate rather than a single anecdote, a rate on its own understates how uncertain that number is. Evan Miller’s statistical treatment of language-model evaluation makes an argument borrowed from experimental science generally: an evaluation score is an estimate drawn from a finite, noisy sample, not a fact about the system, and it should be reported the way any other science reports a measurement — with an uncertainty attached [ 4 ] . His worked illustration is a useful gut check on how much data “enough” actually requires. To reliably detect an absolute difference of three percentage points between two systems, with conventional statistical standards for false positives and false negatives, the calculation runs as follows. Plugging in a 5% false-positive rate, 80% power, and a three-point difference gives roughly 969 independent questions — not thirty, not one hundred, close to a thousand [ 4 ] . Most published agent evaluations, run against benchmarks with a few dozen to a few hundred tasks because each task requires an expensive, stateful, sometimes hours-long rollout, fall well short of that bar by construction, which means many of the differences reported between agent versions or between competing systems are, on this analysis, indistinguishable from noise even when they are reported as a clean percentage-point win. Miller’s own demonstration on real evaluation data found gaps of 3.1 and 2.7 percentage points on two separate benchmarks that did not clear statistical significance once treated properly, and he further showed that once repeated comparisons on the same fixed question set are accounted for with clustered standard errors, the true uncertainty can run more than three times larger than the naive calculation suggests [ 4 ] . His practical recommendation is unglamorous: report the standard error next to the mean, the way other sciences do, rather than a bare percentage that invites over-reading.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Measuring AI Agent Reliability: What the Evidence Actually Supports