Equation 7 · Building a Custom Evaluation Suite for a Production Agent
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol z_β
z_β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives 1.96 and 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly…
Read the full surrounding passage
Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives 1.96 and 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly all of them — have exactly two disciplined options: accept a larger before declaring a regression, watching for a trend across several consecutive merges rather than a single run, or concentrate the expensive repeated-trial budget on the smaller set of tasks where a regression would be most costly and accept thinner evidence everywhere else. Either is defensible; treating a two-point pass-rate wobble on a twenty-task suite as a finding is not.
Sources cited in the article section
- [1] Demystifying evals for AI agents ↗
- [6] Evaluate systematically ↗
- [2] Your AI Product Needs Evals ↗
- [5] Evaluation best practices ↗
- [7] How to build continuous evaluation for AI agents with trace classifications ↗
These citations give research context. Read each source to check which claims it supports.
Return to Building a Custom Evaluation Suite for a Production Agent