Equation 15 · Building a Custom Evaluation Suite for a Production Agent
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol B_day
ay is part of the quantity the equation computes from the expression on the right.
Symbol k_1
is one of the signed contributions combined to compute the quantity on the left.
Symbol n_1
is one of the signed contributions combined to compute the quantity on the left.
Symbol c_1
is one of the signed contributions combined to compute the quantity on the left.
Symbol k_2
is one of the signed contributions combined to compute the quantity on the left.
Symbol n_2
is one of the signed contributions combined to compute the quantity on the left.
Symbol c_2
is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The practical answer most teams converge on is a tiered budget, and it is worth writing the tiering down as an explicit equation rather than an implicit habit, because an implicit habit is what quietly erodes into “we stopped running the full suite” a few months in. If every merge triggers a cheap smoke tier of tasks at unit cost , run times a day, and a slower, more thorough tier of tasks at unit cost , run only times a day, the daily evaluation bill is simply . The first term is the per-merge tax — it has to stay small enough that engineers do not start skipping it, which is the single most common way a CI gate quietly stops functioning. The…
Read the full surrounding passage
The practical answer most teams converge on is a tiered budget, and it is worth writing the tiering down as an explicit equation rather than an implicit habit, because an implicit habit is what quietly erodes into “we stopped running the full suite” a few months in. If every merge triggers a cheap smoke tier of tasks at unit cost , run times a day, and a slower, more thorough tier of tasks at unit cost , run only times a day, the daily evaluation bill is simply . The first term is the per-merge tax — it has to stay small enough that engineers do not start skipping it, which is the single most common way a CI gate quietly stops functioning. The second term is the trust budget — infrequent, thorough, and the number that should scale up before a release rather than on every keystroke. Braintrust’s continuous-evaluation documentation reports exactly this kind of tiering in its online-scoring guidance: score at somewhere between one and ten percent of traffic for high-volume applications, but score one hundred percent of traffic for flows judged critical enough that a missed regression is unacceptable [ 7 ] . That sampling rate is itself a budget decision made the same way — a cheap, wide net most of the time, and a full, expensive check only where the cost of missing something is highest. Writing the split down as two explicit numbers, rather than letting it happen ad hoc, is what keeps the fast tier fast enough to actually run on every change and the slow tier honest enough to actually catch what the fast tier cannot.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Building a Custom Evaluation Suite for a Production Agent