Equation 12 · Why Cost and Latency Belong in the Evaluation Score, Not a Footnote
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol A
at least as good on every axis and strictly better on at least one: [displayed formula].
Symbol B
B is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol s_A
the success indicator or rate with subscript A (at least as good on every axis and strictly better on at least one: [displayed formula]).
Symbol s_B
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol c_A
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol c_B
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol l_A
the expected latency with subscript A (at least as good on every axis and strictly better on at least one: [displayed formula]).
Symbol l_B
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Given that, the more defensible comparison between two agents is not a rank at all. It is a dominance relation. Say agent A Pareto-dominates agent B if A is at least as good on every axis and strictly better on at least one: . The Pareto frontier across a set of candidate agents is the subset that no other candidate dominates — the agents for which improving on any one axis would require giving something up on another. This relation makes no assumption about how much a point of accuracy is worth in dollars or seconds. It only says when one agent is unambiguously not worse than another. That is a weaker claim than a rank, and it is weaker on purpose: it is the claim the…
Read the full surrounding passage
Given that, the more defensible comparison between two agents is not a rank at all. It is a dominance relation. Say agent A Pareto-dominates agent B if A is at least as good on every axis and strictly better on at least one: . The Pareto frontier across a set of candidate agents is the subset that no other candidate dominates — the agents for which improving on any one axis would require giving something up on another. This relation makes no assumption about how much a point of accuracy is worth in dollars or seconds. It only says when one agent is unambiguously not worse than another. That is a weaker claim than a rank, and it is weaker on purpose: it is the claim the data alone actually supports, before anyone’s judgment about relative value has been added to it.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to Why Cost and Latency Belong in the Evaluation Score, Not a Footnote