← Back to article

Equation 20 · How Mechanistic Interpretability Research Is Actually Done

What does this equation mean?

Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt)).\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big).

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cc

Symbol c

c is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

mm

Symbol m

m is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

MM

Symbol M

M is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

xcorruptx_{\mathrm{corrupt}}

Symbol x_corrupt

xcx_corrupt is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

aca_c

Symbol a_c

aca_c is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

accleana_c^{\mathrm{clean}}

Symbol a_c^clean

acca_c^clean is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…
Read the full surrounding passage
The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how much of the clean behaviour component c can restore on its own. The complementary direction, a corrupted value spliced into an otherwise clean run, is noising, and it measures how much clean behaviour depends on c specifically. The two ask different questions, and a careful study reports both, together with the exact corruption and the exact metric used, because Zhang and Nanda showed that changing either of those methodological choices can change which components a patching study identifies as important for the same behaviour in the same model — a localisation claim, they establish, is a claim relative to a stated protocol, not a free-standing fact about the network [ 11 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How Mechanistic Interpretability Research Is Actually Done

See this formula across 1 published context →

Browse the mathematical compendium →