← All parts of this equation

Equation 20 · Part 3 · How Mechanistic Interpretability Research Is Actually Done

Symbol M

Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt)).\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big).
MM

What this part means

M is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

M is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: Patch(c)=m(M(xcorrupt; ac←acclean))−m(M(xcorrupt))\mathrm{Patch}(c) = m\Big(M\big(x_{\mathrm{corrupt}};\ a_c \leftarrow a_c^{\mathrm{clean}}\big)\Big) - m\big(M(x_{\mathrm{corrupt}})\big). This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how much of the clean behaviour component c can restore on…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.