Symbol k
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · Model Evaluation
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…
Equation guide → · Article →Equation 9 · Model Evaluation
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…
Equation guide → · Article →Equation 10 · Model Evaluation
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…
Equation guide → · Article →Equation 20 · Model Evaluation
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
One. Major agentic benchmark leaderboards will begin reporting a repeated-trial statistic (a pass ^k -style figure or equivalent) alongside single-attempt accuracy as a default, not an optional addendum. Disconfirmed if the leading agentic leaderboards in 2030 still report only single-attempt success rates with no standard reliability figure alongside them.
Equation guide → · Article →Equation 19 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
The tau-bench numbers make the distinction concrete rather than abstract. Testing function-calling agents built on frontier models against realistic retail and airline customer-service scenarios — each requiring the agent to follow domain policy, call the right APIs, and leave the backend database in a state matching an annotated goal — the authors found that even strong agents succeeded on well under half of tasks at pass@1, and that reliability across repeats fell sharply as k increased: an agent’s pass ^k score in the retail domain dropped below 25% by k=8 , and airline-domain performance, already lower at k=1 , degraded further from there [ 3 ] . Put in words rather than symbols: an…
Equation guide → · Article →Equation 24 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
State which question is being answered. A best-of- k number (pass@k) and an every-time number (pass ^k ) are both legitimate, and they can diverge sharply on the same system and the same task set [ 3 ] . Reporting one while a reader assumes the other is a common, avoidable source of overconfidence.
Equation guide → · Article →Equation 26 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
One. Vendor system cards and agent-benchmark leaderboards will increasingly report an explicit reliability statistic — a pass ^k -style figure or a confidence interval — alongside a headline pass@1 or pass@k number, because the gap between the two is now well documented rather than speculative. Disconfirmed if major leaderboards in 2029 still report a single unqualified success rate with no repeated-trial or uncertainty disclosure.
Equation guide → · Article →Equation 27 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
A single successful run is evidence that a task is solvable, not evidence that a system solves it reliably, and the statistics that separate the two claims already exist: pass@k for whether at least one of several attempts succeeds, pass ^k for the much stricter question of whether every one of them does, confidence intervals for whether an observed difference is more than noise, time-horizon curves for how reliability changes as a task grows, and a clear accounting of how a number was elicited before it is compared against how a system actually behaves once deployed. None of this is a call for more skepticism in the abstract. It is a call to ask, of any reported agent capability, which of…
Equation guide → · Article →