Symbol k
k appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →Published equation contexts
tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a pas metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption pas makes, and the empty-response exploit breaks exactly it: . averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…
k appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →t is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →i appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 8 · Model Evaluation
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
tau-bench itself — the benchmark in which that first exploit was found — was built to move past shallow grading, simulating a multi-turn conversation between a user (played by a language model) and a tool-using agent, then scoring the conversation against the resulting database state, with a pas metric meant to capture reliability across repeated trials rather than a single lucky success [ 7 ] . The mechanism is worth stating precisely, because it is a real assumption pas makes, and the empty-response exploit breaks exactly it: . averaged over tasks to produce the benchmark’s headline number. The metric is designed to punish an agent whose competence is real but…
Equation guide → · Article →