Symbol hat p
hat p is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
hat p is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · Model Evaluation
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
Equation guide → · Article →