← Mathematical compendium

Published equation contexts

s^i,b=gb(θi)+δi,b+εi,b\hat{s}_{i,b} = g_b(\theta_i) + \delta_{i,b} + \varepsilon_{i,b}

Why this formula appears here

The underlying reason none of these approaches resolves the problem is structural. A published score for model i on benchmark b can be decomposed, at least conceptually, as s^i,b=gb(θi)+δi,b+εi,b\hat{s}_{i,b} = g_b(\theta_i) + \delta_{i,b} + \varepsilon_{i,b}. where θi\theta_i is some unobservable underlying ability vector for model i , gbg_b is the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks, δi,b\delta_{i,b} is a contamination or leakage term specific to that model-benchmark pair, and εi,b\varepsilon_{i,b} is sampling and decoding noise. Because gbg_b differs across benchmarks by construction — a multiple-choice knowledge test and a pairwise human-preference vote are not measuring the same projection of…

Read the full article-specific guide →

Read the representative guide

s^i,b\hat{s}_{i,b}

Symbol hats_i,b

hatsis_i,b is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
gbg_b

Symbol g_b

the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

s^i,b=gb(θi)+δi,b+εi,b,\hat{s}_{i,b} = g_b(\theta_i) + \delta_{i,b} + \varepsilon_{i,b},

Equation 3 · Model Evaluation

The Hardest Unsolved Problems in Frontier AI Model Comparisons

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The underlying reason none of these approaches resolves the problem is structural. A published score for model i on benchmark b can be decomposed, at least conceptually, as s^i,b=gb(θi)+δi,b+εi,b\hat{s}_{i,b} = g_b(\theta_i) + \delta_{i,b} + \varepsilon_{i,b}. where θi\theta_i is some unobservable underlying ability vector for model i , gbg_b is the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks, δi,b\delta_{i,b} is a contamination or leakage term specific to that model-benchmark pair, and εi,b\varepsilon_{i,b} is sampling and decoding noise. Because gbg_b differs across benchmarks by construction — a multiple-choice knowledge test and a pairwise human-preference vote are not measuring the same projection of…

Meanings in this article

Equation guide → · Article →