← Back to article

Equation 22 · Why Cost and Latency Belong in the Evaluation Score, Not a Footnote

What does this equation mean?

λc\lambda_c

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

λc\lambda_c

Symbol lambda_c

lambdaca_c is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

This is exactly why the more careful independent evaluators have converged on publishing the raw triple rather than a single blended figure. Vals AI, an independent third-party evaluator that builds benchmarks for professional domains such as law, tax, and finance, states its position on this directly: “benchmarks often report only accuracy numbers; however, it is important to consider factors such as efficiency, cost, time taken per test, failure modes, and more” [ 4 ] . Its published results report accuracy, latency, and cost as three separate figures per model per benchmark, alongside error bars meant to reflect statistical uncertainty in each one [ 4 ] — leaving the weighting between…
Read the full surrounding passage
This is exactly why the more careful independent evaluators have converged on publishing the raw triple rather than a single blended figure. Vals AI, an independent third-party evaluator that builds benchmarks for professional domains such as law, tax, and finance, states its position on this directly: “benchmarks often report only accuracy numbers; however, it is important to consider factors such as efficiency, cost, time taken per test, failure modes, and more” [ 4 ] . Its published results report accuracy, latency, and cost as three separate figures per model per benchmark, alongside error bars meant to reflect statistical uncertainty in each one [ 4 ] — leaving the weighting between them to the reader rather than baking a λc\lambda_c and λl\lambda_l into the ranking itself. That is a defensible editorial choice, and it is a minority one; most agent benchmarks still publish only s .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Why Cost and Latency Belong in the Evaluation Score, Not a Footnote

Browse the mathematical compendium →