← Back to article

Equation 7 · Building a Custom Evaluation Suite for a Production Agent

What does this equation mean?

zβ≈0.84z_{\beta} \approx 0.84

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

zβz_{\beta}

Symbol z_β

z_β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives zα/2z_{\alpha/2} ≈\approx 1.96 and zβz_{\beta} ≈\approx 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly…
Read the full surrounding passage
Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives zα/2z_{\alpha/2} ≈\approx 1.96 and zβz_{\beta} ≈\approx 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly all of them — have exactly two disciplined options: accept a larger δ\delta before declaring a regression, watching for a trend across several consecutive merges rather than a single run, or concentrate the expensive repeated-trial budget on the smaller set of tasks where a regression would be most costly and accept thinner evidence everywhere else. Either is defensible; treating a two-point pass-rate wobble on a twenty-task suite as a finding is not.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Building a Custom Evaluation Suite for a Production Agent

See this formula across 1 published context →

Browse the mathematical compendium →