Equation 3 · Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol M
M is part of the quantity the equation computes from the expression on the right.
Symbol varepsilon
varepsilon is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
1. Goodhart’s-law metric gaming. The general phenomenon predates language models by decades. Economist Charles Goodhart’s original observation about monetary policy was given its now-standard paraphrase by anthropologist Marilyn Strathern in a 1997 study of Britain’s university audit system: “when a measure becomes a target, it ceases to be a good measure” [ 2 ] . Strathern’s paper is not about machine learning at all — it studies how academic departments reshaped their behavior specifically to satisfy the metrics a national research assessment used to rank them — but the mechanism it documents is exactly the one that recurs in agent evaluation: once a proxy is known and rewarded, effort…
Read the full surrounding passage
1. Goodhart’s-law metric gaming. The general phenomenon predates language models by decades. Economist Charles Goodhart’s original observation about monetary policy was given its now-standard paraphrase by anthropologist Marilyn Strathern in a 1997 study of Britain’s university audit system: “when a measure becomes a target, it ceases to be a good measure” [ 2 ] . Strathern’s paper is not about machine learning at all — it studies how academic departments reshaped their behavior specifically to satisfy the metrics a national research assessment used to rank them — but the mechanism it documents is exactly the one that recurs in agent evaluation: once a proxy is known and rewarded, effort redirects from the underlying goal to the proxy itself. David Manheim and Scott Garrabrant later gave the phenomenon a formal taxonomy, distinguishing at least four distinct failure mechanisms grouped under Goodhart’s name rather than one [ 1 ] . The simplest to state is regressional: if a proxy metric M relates to the true target U by . for some error term uncorrelated with U , then selecting the run with the highest M selects simultaneously for high U and for large favorable — the optimizer is drawn toward the tail of the error term as reliably as toward genuine capability, and the more intensely M is optimized, the larger a share of the observed gain is rather than U .
Sources cited in the surrounding passage
- [2] 'Improving ratings': audit in the British University system ↗
- [1] Categorizing Variants of Goodhart's Law ↗
These citations give research context. Read each source to check which claims it supports.
Return to Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly