← Back to article

Equation 4 · How to Actually Compare Frontier AI Models Without Building a Misleading Leaderboard

What does this equation mean?

θ\theta

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the function of the weights. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

θ\theta

Symbol θ

the function of the weights.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

a function of the weights θ\theta , the harness H\mathcal{H} , the sampling temperature T , the number of samples k and how they are aggregated, the prompt template P\mathcal{P} , the date t a snapshot was queried, plus sampling noise ε\varepsilon . A citation that fixes θ\theta and leaves the other five free has not specified a comparable quantity. OpenAI’s own simple-evals table makes the same point from a different angle: rows the lab ran itself under a disclosed zero-shot chain-of-thought protocol sit beside rows for competitors’ models labelled “unknown” prompt style, sourced from those vendors’ own announcements rather than reproduced on OpenAI’s harness [ 5 ] . Even holding the harness…
Read the full surrounding passage
a function of the weights θ\theta , the harness H\mathcal{H} , the sampling temperature T , the number of samples k and how they are aggregated, the prompt template P\mathcal{P} , the date t a snapshot was queried, plus sampling noise ε\varepsilon . A citation that fixes θ\theta and leaves the other five free has not specified a comparable quantity. OpenAI’s own simple-evals table makes the same point from a different angle: rows the lab ran itself under a disclosed zero-shot chain-of-thought protocol sit beside rows for competitors’ models labelled “unknown” prompt style, sourced from those vendors’ own announcements rather than reproduced on OpenAI’s harness [ 5 ] . Even holding the harness fixed, dated snapshots of the same public model name diverge: that table lists three different GPT-4o release dates scoring 46.0%, 49.9%, and 53.1% on GPQA under the identical simple-evals protocol — a seven-point range attributable to nothing but which week the model was queried [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How to Actually Compare Frontier AI Models Without Building a Misleading Leaderboard

Browse the mathematical compendium →