Equation 6 · Part 1 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
Symbol m
What this part means
m is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
m is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol m→Article meaning
The passage around this formula
The standard instrument is activation patching : run the model on a clean input, run it on a corrupted one, then substitute the activations of a chosen component from one run into the other and measure the change in some behavioural metric. Formally, for a set of components and a metric m ,
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [14] Locating and Editing Factual Associations in GPT ↗
- [17] Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability ↗
- [15] Towards Best Practices of Activation Patching in Language Models: Metrics and Methods ↗
- [16] Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching ↗
These citations provide research context; check each source for the exact claim it supports.