← All parts of this equation

Equation 7 · Part 1 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability

Symbol Δ

Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean).\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right).
Δ\Delta

What this part means

Δ is part of the quantity the equation computes from the expression on the right.

Its job in the formula

Δ is part of the quantity the equation computes from the expression on the right.

The passage around this formula

The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components C\mathcal{C} and a metric m , Δ(C)=m ⁣(Mclean with C set to aCcorrupt)−m ⁣(Mclean)\Delta(\mathcal{C}) = m\!\left(M_{\mathrm{clean}} \text{ with } \mathcal{C} \text{ set to } a_{\mathcal{C}}^{\mathrm{corrupt}}\right) - m\!\left(M_{\mathrm{clean}}\right). Meng and colleagues used a version of this — causal tracing — to localise factual recall, finding that a distinct set of steps in middle-layer feed-forward modules mediate factual predictions at the subject token, and then used the localisation to perform rank-one weight edits that changed specific facts [ 14 ] . The edit is the strongest form of the argument: a…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.