← Back to article

Equation 22 · Measuring AI Agent Reliability: What the Evidence Actually Supports

What does this equation mean?

n  ≳  (zα/2+zβ)29 δ2n \;\gtrsim\; \frac{\left(z_{\alpha/2} + z_{\beta}\right)^{2}}{9\,\delta^{2}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

zβz_{\beta}

Symbol z_β

z_β occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

δ2\delta^{2}

Symbol delta^2

delta2a^2 occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
(zα/2+zβ)2\left(z_{\alpha/2} + z_{\beta}\right)^{2}

Numerator: (z_α/2 + z_β)^2

The complete quantity above the fraction bar.

Understand this part →

See an illustrated explanation →
9 δ29\,\delta^{2}

Denominator: 9delta^2

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

What the article says around this equation

Even when a team does report a rate rather than a single anecdote, a rate on its own understates how uncertain that number is. Evan Miller’s statistical treatment of language-model evaluation makes an argument borrowed from experimental science generally: an evaluation score is an estimate drawn from a finite, noisy sample, not a fact about the system, and it should be reported the way any other science reports a measurement — with an uncertainty attached [ 4 ] . His worked illustration is a useful gut check on how much data “enough” actually requires. To reliably detect an absolute difference of three percentage points between two systems, with conventional statistical standards for false…
Read the full surrounding passage
Even when a team does report a rate rather than a single anecdote, a rate on its own understates how uncertain that number is. Evan Miller’s statistical treatment of language-model evaluation makes an argument borrowed from experimental science generally: an evaluation score is an estimate drawn from a finite, noisy sample, not a fact about the system, and it should be reported the way any other science reports a measurement — with an uncertainty attached [ 4 ] . His worked illustration is a useful gut check on how much data “enough” actually requires. To reliably detect an absolute difference of three percentage points between two systems, with conventional statistical standards for false positives and false negatives, the calculation runs as follows. Plugging in a 5% false-positive rate, 80% power, and a three-point difference gives roughly 969 independent questions — not thirty, not one hundred, close to a thousand [ 4 ] . Most published agent evaluations, run against benchmarks with a few dozen to a few hundred tasks because each task requires an expensive, stateful, sometimes hours-long rollout, fall well short of that bar by construction, which means many of the differences reported between agent versions or between competing systems are, on this analysis, indistinguishable from noise even when they are reported as a clean percentage-point win. Miller’s own demonstration on real evaluation data found gaps of 3.1 and 2.7 percentage points on two separate benchmarks that did not clear statistical significance once treated properly, and he further showed that once repeated comparisons on the same fixed question set are accounted for with clustered standard errors, the true uncertainty can run more than three times larger than the naive calculation suggests [ 4 ] . His practical recommendation is unglamorous: report the standard error next to the mean, the way other sciences do, rather than a bare percentage that invites over-reading.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Measuring AI Agent Reliability: What the Evidence Actually Supports

See this formula across 1 published context →

Browse the mathematical compendium →