Published equation contexts
Why this formula appears here
The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: . where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Read the representative guide
Symbol m_patch
that same metric after the activations of component set are copied from one run into the other.
Read this term in its guide →Symbol m_corrupt
a chosen behavioural metric measured on the clean and corrupted runs.
Read this term in its guide →Symbol m_clean
a chosen behavioural metric measured on the clean and corrupted runs.
Read this term in its guide →Numerator: m_patch(mathcal C) - m_corrupt
The complete quantity above the fraction bar.
Read this term in its guide →Denominator: m_clean - m_corrupt
The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 17 · AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: . where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Meanings in this article
- : copied from one run into the other.
- : that same metric after the activations of component set are copied from one run into the other.
- : a chosen behavioural metric measured on the clean and corrupted runs.
- : a chosen behavioural metric measured on the clean and corrupted runs.