← Mathematical compendium

Published equation contexts

zα/2≈1.96z_{\alpha/2} \approx 1.96

Why this formula appears here

Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives zα/2z_{\alpha/2} ≈\approx 1.96 and zβz_{\beta} ≈\approx 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly…

Read the full article-specific guide →

Read the representative guide

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (2)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

zα/2≈1.96z_{\alpha/2} \approx 1.96

Equation 6 · Model Evaluation

Building a Custom Evaluation Suite for a Production Agent

This equation gives an approximation: it relates the quantities while allowing an approximation.

Plugging in a fairly ordinary case — a baseline pass rate of 70 percent, a minimum shift worth caring about of 10 percentage points, a 5 percent significance level, and 80 percent power — gives zα/2z_{\alpha/2} ≈\approx 1.96 and zβz_{\beta} ≈\approx 0.84 , and the formula returns roughly 330 trials per configuration. That number is the whole argument for treating a single failed run, or even three failed runs out of five, with real suspicion rather than as proof of a regression: at ordinary sample sizes, most of what looks like a pass-rate change between two prompt versions is sampling noise, not signal. Teams that cannot afford several hundred repeated trials on every task for every change — nearly…

Equation guide → · Article →
zα/2≈1.96z_{\alpha/2} \approx 1.96

Equation 6 · Model Evaluation

A Practitioner's Map of Agent Evaluation Frameworks

This equation gives an approximation: it relates the quantities while allowing an approximation.

gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…

Equation guide → · Article →