Symbol T_50
the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t.
Read this term in its guide →Published equation contexts
where is the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t . This is an empirical regression over six years of a specific evaluation methodology, not a law of nature, and METR’s own reporting is explicit that the trend could bend in either direction; it is included here because it is the clearest available answer to “how would you know if agentic coding tools were actually getting more autonomous,” as distinct from “how would you know if a vendor said so.”
the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 2 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where is the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t . This is an empirical regression over six years of a specific evaluation methodology, not a law of nature, and METR’s own reporting is explicit that the trend could bend in either direction; it is included here because it is the clearest available answer to “how would you know if agentic coding tools were actually getting more autonomous,” as distinct from “how would you know if a vendor said so.”
Equation 1 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
It is intuitive to expect an agentic session to improve as it accumulates context about a task. The measured pattern runs the other way. METR’s time-horizon research tracks the length of task — measured in how long a skilled human would take — that a frontier model can complete autonomously with a given reliability, and finds current frontier models near 100% success on tasks taking a human under four minutes, collapsing to under 10% success on tasks taking more than about four hours, with Claude 3.7 Sonnet reported specifically among the models exhibiting this drop-off [ 13 ] . The organization’s headline trend — the task length completable at 50% reliability doubling roughly every seven…