← All parts of this equation

Equation 17 · Part 7 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

subtraction

score(C)=mpatch(C)−mcorruptmclean−mcorrupt,\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}},
subtraction

What this part means

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Its job in the formula

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

The passage around this formula

The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: score(C)=mpatch(C)−mcorruptmclean−mcorrupt\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}}. where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.