← All parts of this equation

Equation 3 · Part 11 · How to Evaluate Codex Beyond SWE-bench and Vendor Scores

Ending index or upper bound: n

p^=1n∑i=1nYi\hat p=\frac{1}{n}\sum_{i=1}^{n}Y_i
nn

What this part means

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Its job in the formula

n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Let YiY_i∈\in\{0,1\} indicate whether evaluation task i is resolved. The familiar score p^=1n∑i=1nYi\hat p=\frac{1}{n}\sum_{i=1}^{n}Y_i. is an estimate of success probability only for a distribution represented by those n instances and the evaluated system. The system is not just a model. It includes context retrieval, tools, prompts, retries, execution limits, and candidate selection. The evaluation distribution is not “software engineering.” It is a constructed sample with inclusion criteria, repository mix, languages, issue styles, and executable tests.

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

See this notation across published equations →

The article lists its research sources here.