Symbol T_50
the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t.
Read this term in its guide →Published equation contexts
A second, independently run measurement effort took a different approach: instead of asking whether an agent could solve one fixed set of tasks, it asked how long a task an agent could complete autonomously with 50% reliability, using human professional completion time as the unit. METR’s March 2025 study, compiled from 170 tasks and more than 800 human timing baselines across software engineering, cybersecurity and general reasoning work, reported that this “50%-task-completion time horizon” for frontier agents had been doubling approximately every seven months for six years running, and that the finding was robust to an order-of-magnitude error in the underlying measurements, which would…
the length of task, measured in human-hours, that a frontier agent could complete autonomously with 50% reliability at time t.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →τ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · AI Agents & Systems
This equation gives an approximation: it relates the quantities while allowing an approximation.
A second, independently run measurement effort took a different approach: instead of asking whether an agent could solve one fixed set of tasks, it asked how long a task an agent could complete autonomously with 50% reliability, using human professional completion time as the unit. METR’s March 2025 study, compiled from 170 tasks and more than 800 human timing baselines across software engineering, cybersecurity and general reasoning work, reported that this “50%-task-completion time horizon” for frontier agents had been doubling approximately every seven months for six years running, and that the finding was robust to an order-of-magnitude error in the underlying measurements, which would…