Equation 2 · What Claude's Capability Evaluations Actually Measure
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol p
p occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol t
t occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol a
a occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol b
b occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →Denominator: 1 + exp(a + b log t)
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The method is a curve fit, not a lookup table, and stating it plainly exposes the assumption it rests on. For a task of human-expert duration t , METR fits a logistic curve to the model’s observed success rate across many tasks of varying length: . with b > 0 , so predicted success falls as human-equivalent task length grows. The reported time horizon at success level x is then the duration at which the fitted curve crosses that threshold, i.e. the t solving p(t) = x . Two things follow directly from this formulation that a single reported number obscures: the curve, not the crossing point, is the actual result, and a different chosen threshold x produces a different…
Read the full surrounding passage
The method is a curve fit, not a lookup table, and stating it plainly exposes the assumption it rests on. For a task of human-expert duration t , METR fits a logistic curve to the model’s observed success rate across many tasks of varying length: . with b > 0 , so predicted success falls as human-equivalent task length grows. The reported time horizon at success level x is then the duration at which the fitted curve crosses that threshold, i.e. the t solving p(t) = x . Two things follow directly from this formulation that a single reported number obscures: the curve, not the crossing point, is the actual result, and a different chosen threshold x produces a different headline number from the identical underlying data.
Sources cited in the article section
- [7] Measuring AI Ability to Complete Long Tasks ↗
- [8] Task-Completion Time Horizons of Frontier AI Models ↗
These citations give research context. Read each source to check which claims it supports.
Return to What Claude's Capability Evaluations Actually Measure