Symbol k
one of the few published metrics designed to recover it, and its scarcity elsewhere in this literature is itself a gap in the evidence.
Read this term in its guide →Published equation contexts
At p = 0.60 and k = 8 , that model predicts roughly 1.7%. The reported figure — under 25% — sits well above that naive prediction, and the direction of the gap is informative on its own: it is only possible if outcomes are not independent draws from one fixed probability, but rather reflect a task population that splits into instances the agent reliably solves and instances it reliably does not, with the reported 60% average blending the two. The practical consequence is that a single success-rate figure understates how often a system that “usually works” will keep failing on the same class of request every single time it is asked, and overstates how often a system that “usually fails” might…
one of the few published metrics designed to recover it, and its scarcity elsewhere in this literature is itself a gap in the evidence.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · AI Infrastructure
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
At p = 0.60 and k = 8 , that model predicts roughly 1.7%. The reported figure — under 25% — sits well above that naive prediction, and the direction of the gap is informative on its own: it is only possible if outcomes are not independent draws from one fixed probability, but rather reflect a task population that splits into instances the agent reliably solves and instances it reliably does not, with the reported 60% average blending the two. The practical consequence is that a single success-rate figure understates how often a system that “usually works” will keep failing on the same class of request every single time it is asked, and overstates how often a system that “usually fails” might…
Equation 20 · AI Agents & Systems
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The tau-bench numbers make the distinction concrete rather than abstract. Testing function-calling agents built on frontier models against realistic retail and airline customer-service scenarios — each requiring the agent to follow domain policy, call the right APIs, and leave the backend database in a state matching an annotated goal — the authors found that even strong agents succeeded on well under half of tasks at pass@1, and that reliability across repeats fell sharply as k increased: an agent’s pass ^k score in the retail domain dropped below 25% by k=8 , and airline-domain performance, already lower at k=1 , degraded further from there [ 3 ] . Put in words rather than symbols: an…
Equation guide → · Article →