← Mathematical compendium

Published equation contexts

score(C)=mpatch(C)−mcorruptmclean−mcorrupt\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}}

Why this formula appears here

The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: score(C)=mpatch(C)−mcorruptmclean−mcorrupt\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}}. where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Read the full article-specific guide →

Read the representative guide

mpatchm_{\text{patch}}

Symbol m_patch

that same metric after the activations of component set C\mathcal C are copied from one run into the other.

Read this term in its guide →
mpatch(C)−mcorruptm_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}

Numerator: m_patch(mathcal C) - m_corrupt

The complete quantity above the fraction bar.

Read this term in its guide →
mclean−mcorruptm_{\text{clean}} - m_{\text{corrupt}}

Denominator: m_clean - m_corrupt

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

score(C)=mpatch(C)−mcorruptmclean−mcorrupt,\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}},

Equation 17 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: score(C)=mpatch(C)−mcorruptmclean−mcorrupt\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}}. where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Meanings in this article

Equation guide → · Article →