← Mathematical compendium

Published equation contexts

k^k

Why this formula appears here

where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…

Read the full article-specific guide →

Read the representative guide

kk

Symbol k

k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (8)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

k^k

Equation 5 · Model Evaluation

The Hardest Unsolved Problems in AI Agent Evaluation and Reliability

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…

Equation guide → · Article →
k^k

Equation 9 · Model Evaluation

The Hardest Unsolved Problems in AI Agent Evaluation and Reliability

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…

Equation guide → · Article →
k^k

Equation 10 · Model Evaluation

The Hardest Unsolved Problems in AI Agent Evaluation and Reliability

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

where pass@ k is the probability that at least one of k attempts succeeds, and pass ^k is the probability that all k succeed [ 3 ] . The two statistics move in opposite directions as k grows: pass@ k climbs toward certainty, which is the right question when a system can retry until something works or a human picks the best of several drafts, while pass ^k falls toward zero, which is the right question for an agent deployed without a human standing by to catch the failures. Yao and colleagues report that state-of-the-art function-calling agents solved under half of τ-bench’s tasks on a single attempt, and that their pass ^k scores fell substantially with repeated trials of the same task under…

Equation guide → · Article →
k^k

Equation 20 · Model Evaluation

The Hardest Unsolved Problems in AI Agent Evaluation and Reliability

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

One. Major agentic benchmark leaderboards will begin reporting a repeated-trial statistic (a pass ^k -style figure or equivalent) alongside single-attempt accuracy as a default, not an optional addendum. Disconfirmed if the leading agentic leaderboards in 2030 still report only single-attempt success rates with no standard reliability figure alongside them.

Equation guide → · Article →
k^k

Equation 19 · AI Agents & Systems

Measuring AI Agent Reliability: What the Evidence Actually Supports

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The tau-bench numbers make the distinction concrete rather than abstract. Testing function-calling agents built on frontier models against realistic retail and airline customer-service scenarios — each requiring the agent to follow domain policy, call the right APIs, and leave the backend database in a state matching an annotated goal — the authors found that even strong agents succeeded on well under half of tasks at pass@1, and that reliability across repeats fell sharply as k increased: an agent’s pass ^k score in the retail domain dropped below 25% by k=8 , and airline-domain performance, already lower at k=1 , degraded further from there [ 3 ] . Put in words rather than symbols: an…

Equation guide → · Article →
k^k

Equation 24 · AI Agents & Systems

Measuring AI Agent Reliability: What the Evidence Actually Supports

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

State which question is being answered. A best-of- k number (pass@k) and an every-time number (pass ^k ) are both legitimate, and they can diverge sharply on the same system and the same task set [ 3 ] . Reporting one while a reader assumes the other is a common, avoidable source of overconfidence.

Equation guide → · Article →
k^k

Equation 26 · AI Agents & Systems

Measuring AI Agent Reliability: What the Evidence Actually Supports

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

One. Vendor system cards and agent-benchmark leaderboards will increasingly report an explicit reliability statistic — a pass ^k -style figure or a confidence interval — alongside a headline pass@1 or pass@k number, because the gap between the two is now well documented rather than speculative. Disconfirmed if major leaderboards in 2029 still report a single unqualified success rate with no repeated-trial or uncertainty disclosure.

Equation guide → · Article →
k^k

Equation 27 · AI Agents & Systems

Measuring AI Agent Reliability: What the Evidence Actually Supports

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

A single successful run is evidence that a task is solvable, not evidence that a system solves it reliably, and the statistics that separate the two claims already exist: pass@k for whether at least one of several attempts succeeds, pass ^k for the much stricter question of whether every one of them does, confidence intervals for whether an observed difference is more than noise, time-horizon curves for how reliability changes as a task grows, and a clear accounting of how a number was elicited before it is compared against how a system actually behaves once deployed. None of this is a call for more skepticism in the abstract. It is a call to ask, of any reported agent capability, which of…

Equation guide → · Article →