← Back to article

Equation 3 · How to Actually Compare Frontier AI Models Without Building a Misleading Leaderboard

What does this equation mean?

s=g(θ,H,T,k,P,t)+ε,s = g(\theta, \mathcal{H}, T, k, \mathcal{P}, t) + \varepsilon,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsg(θ, H, T, k, P, t) + varepsilon
Result or conditions
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ss

Symbol s

the observed score.

Understand this part →

gg

Symbol g

g is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

θ\theta

Symbol θ

the function of the weights.

Understand this part →

H\mathcal{H}

Symbol H

the harness.

Understand this part →

TT

Symbol T

the sampling temperature.

Understand this part →

kk

Symbol k

k is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

P\mathcal{P}

Symbol P

the prompt template.

Understand this part →

tt

Symbol t

t is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

ε\varepsilon

Symbol varepsilon

the plus sampling noise.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Formally, an observed score s is not a property of a model θ\theta alone. It is closer to s=g(θ,H,T,k,P,t)+εs = g(\theta, \mathcal{H}, T, k, \mathcal{P}, t) + \varepsilon. a function of the weights θ\theta , the harness H\mathcal{H} , the sampling temperature T , the number of samples k and how they are aggregated, the prompt template P\mathcal{P} , the date t a snapshot was queried, plus sampling noise ε\varepsilon . A citation that fixes θ\theta and leaves the other five free has not specified a comparable quantity. OpenAI’s own simple-evals table makes the same point from a different angle: rows the lab ran itself under a disclosed zero-shot chain-of-thought protocol sit beside rows for competitors’ models labelled “unknown” prompt style, sourced from…
Read the full surrounding passage
Formally, an observed score s is not a property of a model θ\theta alone. It is closer to s=g(θ,H,T,k,P,t)+εs = g(\theta, \mathcal{H}, T, k, \mathcal{P}, t) + \varepsilon. a function of the weights θ\theta , the harness H\mathcal{H} , the sampling temperature T , the number of samples k and how they are aggregated, the prompt template P\mathcal{P} , the date t a snapshot was queried, plus sampling noise ε\varepsilon . A citation that fixes θ\theta and leaves the other five free has not specified a comparable quantity. OpenAI’s own simple-evals table makes the same point from a different angle: rows the lab ran itself under a disclosed zero-shot chain-of-thought protocol sit beside rows for competitors’ models labelled “unknown” prompt style, sourced from those vendors’ own announcements rather than reproduced on OpenAI’s harness [ 5 ] . Even holding the harness fixed, dated snapshots of the same public model name diverge: that table lists three different GPT-4o release dates scoring 46.0%, 49.9%, and 53.1% on GPQA under the identical simple-evals protocol — a seven-point range attributable to nothing but which week the model was queried [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How to Actually Compare Frontier AI Models Without Building a Misleading Leaderboard

See this formula across 1 published context →

Browse the mathematical compendium →