← All parts of this equation

Equation 7 · Part 1 · Running an Interpretability Investigation That Holds Up

Symbol Δ

Δ(C)=m(Mclean→C←corrupt)−m(Mclean).\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big).
Δ\Delta

What this part means

Δ is part of the quantity the equation computes from the expression on the right.

Its job in the formula

Δ is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Once the question is scoped, the experiment that tests it is a causal one by convention in this field, and the convention exists for a good reason: a component correlated with a behaviour has told you nothing about whether the model uses it, while a component that changes the behaviour when intervened on has told you something. The standard instrument is activation patching — running the model on a clean input, running it again on a corrupted one, then substituting the corrupted run’s activations into the clean run at a chosen set of components C\mathcal{C} and measuring the change in some behavioural metric m : Δ(C)=m(Mclean→C←corrupt)−m(Mclean)\Delta(\mathcal{C}) = m\big(M_{\mathrm{clean} \to \mathcal{C} \leftarrow \mathrm{corrupt}}\big) - m\big(M_{\mathrm{clean}}\big). The choice buried inside that formula — what counts as…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.