← Mathematical compendium

Published equation contexts

R=1−∑i=1nwi⋅1[trial i failed]R = 1 - \sum_{i=1}^{n} w_i \cdot \mathbb{1}[\text{trial } i \text{ failed}]

Why this formula appears here

and contrast it with a severity-weighted version, R=1−∑i=1nwi⋅1[trial i failed]R = 1 - \sum_{i=1}^{n} w_i \cdot \mathbb{1}[\text{trial } i \text{ failed}]. where wiw_i scales each failure by how costly it actually was. Two agents can share an identical p^\hat{p} of, say, ninety-five percent, while one of them fails harmlessly - an unhelpful but reversible answer - and the other fails catastrophically - an irreversible transaction, a deleted repository, a wrong medical dosage recommendation - on that same five percent. No standard agent benchmark publishes R , because assigning a defensible wiw_i requires a judgment about real-world consequence that a replayable, sandboxed task suite is not built to carry, and because the tasks that would carry the highest weights are, not…

Read the full article-specific guide →

Read the representative guide

ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
nn

Symbol n

n appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
i=1i=1

Starting index or lower bound: i=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →
nn

Ending index or upper bound: n

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

R=1−∑i=1nwi⋅1[trial i failed],R = 1 - \sum_{i=1}^{n} w_i \cdot \mathbb{1}[\text{trial } i \text{ failed}],

Equation 15 · Model Evaluation

The Hardest Unsolved Problems in AI Agent Evaluation and Reliability

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

and contrast it with a severity-weighted version, R=1−∑i=1nwi⋅1[trial i failed]R = 1 - \sum_{i=1}^{n} w_i \cdot \mathbb{1}[\text{trial } i \text{ failed}]. where wiw_i scales each failure by how costly it actually was. Two agents can share an identical p^\hat{p} of, say, ninety-five percent, while one of them fails harmlessly - an unhelpful but reversible answer - and the other fails catastrophically - an irreversible transaction, a deleted repository, a wrong medical dosage recommendation - on that same five percent. No standard agent benchmark publishes R , because assigning a defensible wiw_i requires a judgment about real-world consequence that a replayable, sandboxed task suite is not built to carry, and because the tasks that would carry the highest weights are, not…

Meanings in this article

  • RR: the no standard agent benchmark publishes.
Equation guide → · Article →