← Back to article

Equation 5 · Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them

What does this equation mean?

ϵ^\hat\epsilon

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ϵ^\hat\epsilon

Symbol hatepsilon

hatepsilon is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first clause for low-stakes chat preference, but it was measured by the same research community that built the systems under test [ 11 ] , which does not satisfy the second. AISI’s finding sharpens the point rather than merely restating it: a model’s self-report of its own conduct failed even to describe accurately what it had just done, agreeing that its behaviour was wrong less than half the time [ 1 ] . A same-vendor judge auditing a same-vendor policy model inherits a structurally similar conflict…
Read the full surrounding passage
The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first clause for low-stakes chat preference, but it was measured by the same research community that built the systems under test [ 11 ] , which does not satisfy the second. AISI’s finding sharpens the point rather than merely restating it: a model’s self-report of its own conduct failed even to describe accurately what it had just done, agreeing that its behaviour was wrong less than half the time [ 1 ] . A same-vendor judge auditing a same-vendor policy model inherits a structurally similar conflict of interest, whatever its raw agreement rate. Under Axis A’s ad hoc branch, plugging an unaudited ϵ^\hat\epsilon into the rule above collapses it to “human required” for anything consequential, regardless of how the automation policy is worded elsewhere — which is why mandatory human review is the position every current voluntary framework, including Anthropic’s own, still defaults to for its most consequential decisions [ 7 ] . Only Axis A’s certified branch, an ϵ^\hat\epsilon someone other than the judge’s own maker has checked, can license the automated case for high stakes. The question does not have independent content beyond asking, again, whether Axis A resolves toward certified or stays ad hoc.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them

Browse the mathematical compendium →