← Back to article

Equation 5 · Building a Custom Evaluation Suite for a Production Agent

What does this equation mean?

n≈2 p(1−p) (zα/2+zβ)2δ2.n \approx \frac{2\,p(1-p)\,(z_{\alpha/2} + z_{\beta})^2}{\delta^2} .

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

pp

Symbol p

p occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

zβz_{\beta}

Symbol z_β

z_β occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

δ2\delta^2

Symbol delta^2

delta2a^2 occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
2 p(1−p) (zα/2+zβ)22\,p(1-p)\,(z_{\alpha/2} + z_{\beta})^2

Numerator: 2p(1-p)(z_α/2 + z_β)^2

The complete quantity above the fraction bar.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

Because a “pass” verdict is itself often a noisy measurement, not a fact, it is worth being explicit about how much noise a single run’s outcome carries before treating a change in the pass rate as real. If a task’s true pass rate is p under a baseline configuration and a candidate change is worth detecting only once it shifts that rate by at least δ\delta , the number of repeated trials needed per configuration to detect the shift reliably — at significance level α\alpha and statistical power 1-β\beta — is approximately n≈2 p(1−p) (zα/2+zβ)2δ2n \approx \frac{2\,p(1-p)\,(z_{\alpha/2} + z_{\beta})^2}{\delta^2} . Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent…
Read the full surrounding passage
Because a “pass” verdict is itself often a noisy measurement, not a fact, it is worth being explicit about how much noise a single run’s outcome carries before treating a change in the pass rate as real. If a task’s true pass rate is p under a baseline configuration and a candidate change is worth detecting only once it shifts that rate by at least δ\delta , the number of repeated trials needed per configuration to detect the shift reliably — at significance level α\alpha and statistical power 1-β\beta — is approximately n≈2 p(1−p) (zα/2+zβ)2δ2n \approx \frac{2\,p(1-p)\,(z_{\alpha/2} + z_{\beta})^2}{\delta^2} . Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives zα/2z_{\alpha/2} ≈\approx 1.96 and zβz_{\beta} ≈\approx 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly all of them — have exactly two disciplined options: accept a larger δ\delta before declaring a regression, watching for a trend across several consecutive merges rather than a single run, or concentrate the expensive repeated-trial budget on the smaller set of tasks where a regression would be most costly and accept thinner evidence everywhere else. Either is defensible; treating a two-point pass-rate wobble on a twenty-task suite as a finding is not.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Building a Custom Evaluation Suite for a Production Agent

See this formula across 1 published context →

Browse the mathematical compendium →