Equation 8 · Why Cost and Latency Belong in the Evaluation Score, Not a Footnote
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the expected latency. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The trouble with discarding c and l is not that they are minor details. It is that “cost” and “latency” are themselves not single, unambiguous numbers waiting to be read off a meter. Dehghani and colleagues made this point rigorously for model efficiency in general, cataloguing the many different cost indicators researchers use — parameter count, FLOPs, throughput, peak memory, wall-clock time — and showing experimentally that these indicators routinely disagree with one another: an architecture that looks efficient on one indicator looks inefficient on another, and reporting only one or two of them, as is common practice, produces a partial and sometimes actively misleading picture of what…
Read the full surrounding passage
The trouble with discarding c and l is not that they are minor details. It is that “cost” and “latency” are themselves not single, unambiguous numbers waiting to be read off a meter. Dehghani and colleagues made this point rigorously for model efficiency in general, cataloguing the many different cost indicators researchers use — parameter count, FLOPs, throughput, peak memory, wall-clock time — and showing experimentally that these indicators routinely disagree with one another: an architecture that looks efficient on one indicator looks inefficient on another, and reporting only one or two of them, as is common practice, produces a partial and sometimes actively misleading picture of what a model actually costs to run [ 6 ] . An agent adds a further layer on top of that problem, because its cost and latency are not fixed properties of a checkpoint the way parameter count is; they are properties of a whole trajectory of tool calls, retries, and context growth, measured under whatever conditions the evaluator happened to run.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Why Cost and Latency Belong in the Evaluation Score, Not a Footnote