Equation 5 · Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hatepsilon
hatepsilon is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first clause for low-stakes chat preference, but it was measured by the same research community that built the systems under test [ 11 ] , which does not satisfy the second. AISI’s finding sharpens the point rather than merely restating it: a model’s self-report of its own conduct failed even to describe accurately what it had just done, agreeing that its behaviour was wrong less than half the time [ 1 ] . A same-vendor judge auditing a same-vendor policy model inherits a structurally similar conflict…
Read the full surrounding passage
The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first clause for low-stakes chat preference, but it was measured by the same research community that built the systems under test [ 11 ] , which does not satisfy the second. AISI’s finding sharpens the point rather than merely restating it: a model’s self-report of its own conduct failed even to describe accurately what it had just done, agreeing that its behaviour was wrong less than half the time [ 1 ] . A same-vendor judge auditing a same-vendor policy model inherits a structurally similar conflict of interest, whatever its raw agreement rate. Under Axis A’s ad hoc branch, plugging an unaudited into the rule above collapses it to “human required” for anything consequential, regardless of how the automation policy is worded elsewhere — which is why mandatory human review is the position every current voluntary framework, including Anthropic’s own, still defaults to for its most consequential decisions [ 7 ] . Only Axis A’s certified branch, an someone other than the judge’s own maker has checked, can license the automated case for high stakes. The question does not have independent content beyond asking, again, whether Axis A resolves toward certified or stays ad hoc.
Sources cited in the surrounding passage
- [11] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena ↗
- [1] Cheating Behaviour in Frontier Model Evaluations ↗
- [7] Anthropic's Responsible Scaling Policy ↗
These citations give research context. Read each source to check which claims it supports.
Return to Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them