Equation 28 · Why Cost and Latency Belong in the Evaluation Score, Not a Footnote
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L_sequential
equential is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol k
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol barl
barl is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol L_parallel
arallel is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol epsilon
a small scheduling and aggregation overhead that grows slowly with k.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
Start with how an agent handles its own failures. Suppose a task is retried up to k times under two different execution policies: sequential retry, where each attempt waits for the previous one to finish before starting, and parallel sampling, where all k attempts run concurrently and the first success is taken. Cost is roughly indifferent to which policy was used — the compute consumed scales with the number of attempts made, C(k) k , regardless of whether they ran one after another or all at once. Latency is not indifferent at all: . where is a small scheduling and aggregation overhead that grows slowly with k . Two evaluations that both report…
Read the full surrounding passage
Start with how an agent handles its own failures. Suppose a task is retried up to k times under two different execution policies: sequential retry, where each attempt waits for the previous one to finish before starting, and parallel sampling, where all k attempts run concurrently and the first success is taken. Cost is roughly indifferent to which policy was used — the compute consumed scales with the number of attempts made, C(k) k , regardless of whether they ran one after another or all at once. Latency is not indifferent at all: . where is a small scheduling and aggregation overhead that grows slowly with k . Two evaluations that both report “an agent retried up to five times” can therefore report latency figures that differ by nearly a factor of five, purely as an artifact of whether those five attempts ran one after another or side by side, with cost looking nearly identical between them. A latency number published without stating the retry and parallelism policy behind it is not comparable to another latency number published under a different policy, even when both describe the same underlying agent on the same task.
Sources cited in the article section
- [8] Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks ↗
- [10] Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores ↗
These citations give research context. Read each source to check which claims it supports.
Return to Why Cost and Latency Belong in the Evaluation Score, Not a Footnote