Published equation contexts
Why this formula appears here
and contrast it with a severity-weighted version, . where scales each failure by how costly it actually was. Two agents can share an identical of, say, ninety-five percent, while one of them fails harmlessly - an unhelpful but reversible answer - and the other fails catastrophically - an irreversible transaction, a deleted repository, a wrong medical dosage recommendation - on that same five percent. No standard agent benchmark publishes R , because assigning a defensible requires a judgment about real-world consequence that a replayable, sandboxed task suite is not built to carry, and because the tasks that would carry the highest weights are, not…
Read the representative guide
Symbol i
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →Symbol n
n appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →Symbol w_i
is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Starting index or lower bound: i=1
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →Ending index or upper bound: n
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Read this term in its guide →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 15 · Model Evaluation
The Hardest Unsolved Problems in AI Agent Evaluation and Reliability
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
and contrast it with a severity-weighted version, . where scales each failure by how costly it actually was. Two agents can share an identical of, say, ninety-five percent, while one of them fails harmlessly - an unhelpful but reversible answer - and the other fails catastrophically - an irreversible transaction, a deleted repository, a wrong medical dosage recommendation - on that same five percent. No standard agent benchmark publishes R , because assigning a defensible requires a judgment about real-world consequence that a replayable, sandboxed task suite is not built to carry, and because the tasks that would carry the highest weights are, not…