Equation 3 · How to Evaluate Codex Beyond SWE-bench and Vendor Scores
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hat p
hat p is part of the quantity the equation computes from the expression on the right.
Symbol n
n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol Y_i
an estimate of success probability only for a distribution represented by those n instances and the evaluated system.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Starting index or lower bound: i=1
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Ending index or upper bound: n
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Let \{0,1\} indicate whether evaluation task i is resolved. The familiar score . is an estimate of success probability only for a distribution represented by those n instances and the evaluated system. The system is not just a model. It includes context retrieval, tools, prompts, retries, execution limits, and candidate selection. The evaluation distribution is not “software engineering.” It is a constructed sample with inclusion criteria, repository mix, languages, issue styles, and executable tests.
For background, read the article’s source list.
Return to How to Evaluate Codex Beyond SWE-bench and Vendor Scores