← All parts of this equation

Equation 17 · Part 6 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

fraction

score(C)=mpatch(C)−mcorruptmclean−mcorrupt,\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}},
fraction

What this part means

Divide the expression above the line by the one below it.

Its job in the formula

The expression above the fraction bar is divided by the complete expression below it. The denominator must not be zero.

The passage around this formula

The general operation, activation patching, is usually reported as a normalised score rather than a raw difference, isolating exactly how much of the gap between clean and corrupted behaviour a given component set closes: score(C)=mpatch(C)−mcorruptmclean−mcorrupt\text{score}(\mathcal C) = \frac{m_{\text{patch}}(\mathcal C) - m_{\text{corrupt}}}{m_{\text{clean}} - m_{\text{corrupt}}}. where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.