Equation 4 · Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol a
a is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol hatepsilon
hatepsilon appears in the objective or constraint used by the optimization on the right.
Symbol epsilon_max
epsiloax appears in the objective or constraint used by the optimization on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Whether automated LLM-judge evaluation becomes trusted for high-stakes decisions is not a third axis; it is mostly a readout of Axis A applied to one specific evaluator. A useful way to see the dependency is to write the automation decision as a threshold rule. Let be the estimated error rate of a judge on a decision class a , and let be the maximum error a policy is willing to tolerate for that stakes class. A defensible automation rule is . The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first…
Read the full surrounding passage
Whether automated LLM-judge evaluation becomes trusted for high-stakes decisions is not a third axis; it is mostly a readout of Axis A applied to one specific evaluator. A useful way to see the dependency is to write the automation decision as a threshold rule. Let be the estimated error rate of a judge on a decision class a , and let be the maximum error a policy is willing to tolerate for that stakes class. A defensible automation rule is . The rule only licenses automation where both clauses hold, and the second clause is the one Axis A supplies or withholds. Zheng and colleagues’ own eighty-percent figure plausibly satisfies the first clause for low-stakes chat preference, but it was measured by the same research community that built the systems under test [ 11 ] , which does not satisfy the second. AISI’s finding sharpens the point rather than merely restating it: a model’s self-report of its own conduct failed even to describe accurately what it had just done, agreeing that its behaviour was wrong less than half the time [ 1 ] . A same-vendor judge auditing a same-vendor policy model inherits a structurally similar conflict of interest, whatever its raw agreement rate. Under Axis A’s ad hoc branch, plugging an unaudited into the rule above collapses it to “human required” for anything consequential, regardless of how the automation policy is worded elsewhere — which is why mandatory human review is the position every current voluntary framework, including Anthropic’s own, still defaults to for its most consequential decisions [ 7 ] . Only Axis A’s certified branch, an someone other than the judge’s own maker has checked, can license the automated case for high stakes. The question does not have independent content beyond asking, again, whether Axis A resolves toward certified or stays ad hoc.
Sources cited in the surrounding passage
- [11] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena ↗
- [1] Cheating Behaviour in Frontier Model Evaluations ↗
- [7] Anthropic's Responsible Scaling Policy ↗
These citations give research context. Read each source to check which claims it supports.
Return to Agent Evaluation in 2035: Two Axes, Four Scenarios, and What Would Falsify Them