Equation 20 · How Mechanistic Interpretability Research Is Actually Done
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol c
c is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol m
m is one of the signed contributions combined to compute the quantity on the left.
Symbol M
M is one of the signed contributions combined to compute the quantity on the left.
Symbol x_corrupt
orrupt is one of the signed contributions combined to compute the quantity on the left.
Symbol a_c
is one of the signed contributions combined to compute the quantity on the left.
Symbol a_c^clean
lean is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: . This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…
Read the full surrounding passage
The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: . This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how much of the clean behaviour component c can restore on its own. The complementary direction, a corrupted value spliced into an otherwise clean run, is noising, and it measures how much clean behaviour depends on c specifically. The two ask different questions, and a careful study reports both, together with the exact corruption and the exact metric used, because Zhang and Nanda showed that changing either of those methodological choices can change which components a patching study identifies as important for the same behaviour in the same model — a localisation claim, they establish, is a claim relative to a stated protocol, not a free-standing fact about the network [ 11 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How Mechanistic Interpretability Research Is Actually Done