← Mathematical compendium

Published equation contexts

Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big)

Why this formula appears here

The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…

Read the full article-specific guide →

Read the representative guide

xcorruptx_{\mathrm{corrupt}}

Symbol x_corrupt

xcx_corrupt is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
accleana_c^{\mathrm{clean}}

Symbol a_c^clean

acca_c^clean is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt)).\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big).

Equation 20 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…

Equation guide → · Article →