Equation 5 · Building a Custom Evaluation Suite for a Production Agent
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol n
n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol p
p occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol z_α/2
z_α/2 occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol z_β
z_β occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol delta^2
delt occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Numerator: 2p(1-p)(z_α/2 + z_β)^2
The complete quantity above the fraction bar.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
Because a “pass” verdict is itself often a noisy measurement, not a fact, it is worth being explicit about how much noise a single run’s outcome carries before treating a change in the pass rate as real. If a task’s true pass rate is p under a baseline configuration and a candidate change is worth detecting only once it shifts that rate by at least , the number of repeated trials needed per configuration to detect the shift reliably — at significance level and statistical power 1- — is approximately . Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent…
Read the full surrounding passage
Because a “pass” verdict is itself often a noisy measurement, not a fact, it is worth being explicit about how much noise a single run’s outcome carries before treating a change in the pass rate as real. If a task’s true pass rate is p under a baseline configuration and a candidate change is worth detecting only once it shifts that rate by at least , the number of repeated trials needed per configuration to detect the shift reliably — at significance level and statistical power 1- — is approximately . Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives 1.96 and 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly all of them — have exactly two disciplined options: accept a larger before declaring a regression, watching for a trend across several consecutive merges rather than a single run, or concentrate the expensive repeated-trial budget on the smaller set of tasks where a regression would be most costly and accept thinner evidence everywhere else. Either is defensible; treating a two-point pass-rate wobble on a twenty-task suite as a finding is not.
Sources cited in the article section
- [1] Demystifying evals for AI agents ↗
- [6] Evaluate systematically ↗
- [2] Your AI Product Needs Evals ↗
- [5] Evaluation best practices ↗
- [7] How to build continuous evaluation for AI agents with trace classifications ↗
These citations give research context. Read each source to check which claims it supports.
Return to Building a Custom Evaluation Suite for a Production Agent