← All parts of this equation

Equation 4 · Part 1 · Measuring Frontier Models: Contamination, Variance, and What a Score Can Support

Symbol hats

s^=1n∑i=1n1 ⁣[ success on xi ],xi∼Dbench.\hat{s} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\!\left[\,\text{success on } x_i\,\right], \qquad x_i \sim \mathcal{D}_{\mathrm{bench}} .
s^\hat{s}

What this part means

hats is part of the quantity the equation computes from the expression on the right.

Its job in the formula

hats is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Reframe the object. A benchmark score s^\hat{s} is an estimator of a population quantity s — the model’s success rate over some distribution of tasks D\mathcal{D} that somebody hopes resembles the work you actually have. Every property that makes an estimator trustworthy applies: s^=1n∑i=1n1 ⁣[ success on xi ],xi∼Dbench\hat{s} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\!\left[\,\text{success on } x_i\,\right], \qquad x_i \sim \mathcal{D}_{\mathrm{bench}} . The number is only as good as three things: whether Dbench\mathcal{D}_{\mathrm{bench}} resembles D\mathcal{D} , whether the xix_i are genuinely held out, and whether the indicator is measured with enough repetition to characterise its spread. All three fail routinely, and they fail in different directions.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

The article lists its research sources here.