← Mathematical compendium

Published equation contexts

Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right)

Why this formula appears here

The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components C\mathcal{C} and a metric m , Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right). Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a…

Read the full article-specific guide →

Read the representative guide

C\mathcal{C}

Symbol C

C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
McleanM_{\mathrm{clean}}

Symbol M_clean

McM_clean is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
aCcorrupta_{\mathcal{C}}^{\mathrm{corrupt}}

Symbol a_C^corrupt

aCca_C^corrupt is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean).\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right).

Equation 7 · Interpretability

What a Circuit Explains: The State and Limits of Mechanistic Interpretability

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components C\mathcal{C} and a metric m , Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right). Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a…

Equation guide → · Article →