← Mathematical compendium

Published equation contexts

Δ(C)=m(Mclean→C←corrupt)−m(Mclean)\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big)

Why this formula appears here

Once the question is scoped, the experiment that tests it is a causal one by convention in this field, and the convention exists for a good reason: a component correlated with a behaviour has told you nothing about whether the model uses it, while a component that changes the behaviour when intervened on has told you something. The standard instrument is activation patching — running the model on a clean input, running it again on a corrupted one, then substituting the corrupted run’s activations into the clean run at a chosen set of components C\mathcal{C} and measuring the change in some behavioural metric m : Δ(C)=m(Mclean→C←corrupt)−m(Mclean)\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big). The choice buried inside that formula — what counts as…

Read the full article-specific guide →

Read the representative guide

C\mathcal{C}

Symbol C

C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
Mclean→C←corruptM_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}

Symbol M_clean to C arrow corrupt

McM_clean to C arrow corrupt is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
McleanM_{\mathrm{clean}}

Symbol M_clean

McM_clean is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Δ(C)=m(Mclean→C←corrupt)−m(Mclean).\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big).

Equation 7 · AI Research

Running an Interpretability Investigation That Holds Up

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Once the question is scoped, the experiment that tests it is a causal one by convention in this field, and the convention exists for a good reason: a component correlated with a behaviour has told you nothing about whether the model uses it, while a component that changes the behaviour when intervened on has told you something. The standard instrument is activation patching — running the model on a clean input, running it again on a corrupted one, then substituting the corrupted run’s activations into the clean run at a chosen set of components C\mathcal{C} and measuring the change in some behavioural metric m : Δ(C)=m(Mclean→C←corrupt)−m(Mclean)\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big). The choice buried inside that formula — what counts as…

Equation guide → · Article →