Symbol hats_i,b
hat,b is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
The underlying reason none of these approaches resolves the problem is structural. A published score for model i on benchmark b can be decomposed, at least conceptually, as . where is some unobservable underlying ability vector for model i , is the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks, is a contamination or leakage term specific to that model-benchmark pair, and is sampling and decoding noise. Because differs across benchmarks by construction — a multiple-choice knowledge test and a pairwise human-preference vote are not measuring the same projection of…
hat,b is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks.
Read this term in its guide →a contamination or leakage term specific to that model-benchmark pair.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 3 · Model Evaluation
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The underlying reason none of these approaches resolves the problem is structural. A published score for model i on benchmark b can be decomposed, at least conceptually, as . where is some unobservable underlying ability vector for model i , is the mapping from ability to expected score that is specific to benchmark b and not shared across benchmarks, is a contamination or leakage term specific to that model-benchmark pair, and is sampling and decoding noise. Because differs across benchmarks by construction — a multiple-choice knowledge test and a pairwise human-preference vote are not measuring the same projection of…