← All parts of this equation

Equation 4 · Part 11 · Measuring Frontier Models: Contamination, Variance, and What a Score Can Support

Starting index or lower bound: i=1

s^=1n∑i=1n1 ⁣[ success on xi ],xi∼Dbench.\hat{s} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\!\left[\,\text{success on } x_i\,\right], \qquad x_i \sim \mathcal{D}_{\mathrm{bench}} .
i=1i=1

What this part means

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Its job in the formula

i=1 appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

Reframe the object. A benchmark score s^\hat{s} is an estimator of a population quantity s — the model’s success rate over some distribution of tasks D\mathcal{D} that somebody hopes resembles the work you actually have. Every property that makes an estimator trustworthy applies: s^=1n∑i=1n1 ⁣[ success on xi ],xi∼Dbench\hat{s} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\!\left[\,\text{success on } x_i\,\right], \qquad x_i \sim \mathcal{D}_{\mathrm{bench}} . The number is only as good as three things: whether Dbench\mathcal{D}_{\mathrm{bench}} resembles D\mathcal{D} , whether the xix_i are genuinely held out, and whether the indicator is measured with enough repetition to characterise its spread. All three fail routinely, and they fail in different directions.

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

The article lists its research sources here.