Equation 17 · Part 4 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Symbol m_clean
What this part means
a chosen behavioural metric measured on the clean and corrupted runs.
Its job in the formula
lean occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Symbol m_clean→Article meaning
Where the article explains it
where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other.
The passage around this formula
The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: . where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [9] Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias ↗
- [10] Locating and Editing Factual Associations in GPT ↗
- [11] How to Use and Interpret Activation Patching ↗
These citations provide research context; check each source for the exact claim it supports.