← All parts of this equation

Equation 20 · Part 4 · How Mechanistic Interpretability Research Is Actually Done

Symbol x_corrupt

Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt)).\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big).
xcorruptx_{\mathrm{corrupt}}

What this part means

xcx_corrupt is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

xcx_corrupt is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.