Equation 1 · Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol M
M is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
1. Goodhart’s-law metric gaming. The general phenomenon predates language models by decades. Economist Charles Goodhart’s original observation about monetary policy was given its now-standard paraphrase by anthropologist Marilyn Strathern in a 1997 study of Britain’s university audit system: “when a measure becomes a target, it ceases to be a good measure” [ 2 ] . Strathern’s paper is not about machine learning at all — it studies how academic departments reshaped their behavior specifically to satisfy the metrics a national research assessment used to rank them — but the mechanism it documents is exactly the one that recurs in agent evaluation: once a proxy is known and rewarded, effort…
Read the full surrounding passage
1. Goodhart’s-law metric gaming. The general phenomenon predates language models by decades. Economist Charles Goodhart’s original observation about monetary policy was given its now-standard paraphrase by anthropologist Marilyn Strathern in a 1997 study of Britain’s university audit system: “when a measure becomes a target, it ceases to be a good measure” [ 2 ] . Strathern’s paper is not about machine learning at all — it studies how academic departments reshaped their behavior specifically to satisfy the metrics a national research assessment used to rank them — but the mechanism it documents is exactly the one that recurs in agent evaluation: once a proxy is known and rewarded, effort redirects from the underlying goal to the proxy itself. David Manheim and Scott Garrabrant later gave the phenomenon a formal taxonomy, distinguishing at least four distinct failure mechanisms grouped under Goodhart’s name rather than one [ 1 ] . The simplest to state is regressional: if a proxy metric M relates to the true target U by
Sources cited in the surrounding passage
- [2] 'Improving ratings': audit in the British University system ↗
- [1] Categorizing Variants of Goodhart's Law ↗
These citations give research context. Read each source to check which claims it supports.
Return to Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly