← Back to article

Equation 7 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability

What does this equation mean?

Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean).\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right).

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Δ\Delta

Symbol Δ

Δ is part of the quantity the equation computes from the expression on the right.

Understand this part →

C\mathcal{C}

Symbol C

C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

mm

Symbol m

m is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

McleanM_{\mathrm{clean}}

Symbol M_clean

McM_clean is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

aCcorrupta_{\mathcal{C}}^{\mathrm{corrupt}}

Symbol a_C^corrupt

aCca_C^corrupt is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components C\mathcal{C} and a metric m , Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right). Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a…
Read the full surrounding passage
The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components C\mathcal{C} and a metric m , Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right). Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a localisation claim that supports a successful targeted modification has done more than describe.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What a Circuit Explains: The State and Limits of Mechanistic Interpretability

See this formula across 1 published context →

Browse the mathematical compendium →