Symbol T_50
the task duration completable at 50% reliability at year t , METR’s fit is [displayed formula].
Read this term in its guide →Published equation contexts
It is intuitive to expect an agentic session to improve as it accumulates context about a task. The measured pattern runs the other way. METR’s time-horizon research tracks the length of task — measured in how long a skilled human would take — that a frontier model can complete autonomously with a given reliability, and finds current frontier models near 100% success on tasks taking a human under four minutes, collapsing to under 10% success on tasks taking more than about four hours, with Claude 3.7 Sonnet reported specifically among the models exhibiting this drop-off [ 13 ] . The organization’s headline trend — the task length completable at 50% reliability doubling roughly every seven…
the task duration completable at 50% reliability at year t , METR’s fit is [displayed formula].
Read this term in its guide →t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →sensitive to task selection and human-baseline methodology, even though the qualitative exponential trend is one they hold with more confidence [ 13 ].
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 3 · AI Agents & Systems
This equation gives an approximation: it relates the quantities while allowing an approximation.
It is intuitive to expect an agentic session to improve as it accumulates context about a task. The measured pattern runs the other way. METR’s time-horizon research tracks the length of task — measured in how long a skilled human would take — that a frontier model can complete autonomously with a given reliability, and finds current frontier models near 100% success on tasks taking a human under four minutes, collapsing to under 10% success on tasks taking more than about four hours, with Claude 3.7 Sonnet reported specifically among the models exhibiting this drop-off [ 13 ] . The organization’s headline trend — the task length completable at 50% reliability doubling roughly every seven…