← All parts of this equation

Equation 7 · Part 2 · Running an Interpretability Investigation That Holds Up

Symbol C

Δ(C)=m(Mclean→C←corrupt)−m(Mclean).\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big).
C\mathcal{C}

What this part means

C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

…The standard instrument is activation patching — running the model on a clean input, running it again on a corrupted one, then substituting the corrupted run’s activations into the clean run at a chosen set of components C\mathcal{C} and measuring the change in some behavioural metric m : Δ(C)=m(Mclean→C←corrupt)−m(Mclean)\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big). The choice buried inside that formula — what counts as “corrupt” — is not a detail. Zeroing an activation, replacing it with the dataset mean, or resampling it from an unrelated input each encode a different…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.