Equation 7 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol Δ
Δ is part of the quantity the equation computes from the expression on the right.
Symbol C
C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol m
m is one of the signed contributions combined to compute the quantity on the left.
Symbol M_clean
lean is one of the signed contributions combined to compute the quantity on the left.
Symbol a_C^corrupt
orrupt is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components and a metric m , . Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a…
Read the full surrounding passage
The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components and a metric m , . Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a localisation claim that supports a successful targeted modification has done more than describe.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What a Circuit Explains: The State and Limits of Mechanistic Interpretability