← Back to article

Equation 18 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

mcleanm_{\text{clean}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

a chosen behavioural metric measured on the clean and corrupted runs. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

mcleanm_{\text{clean}}

Symbol m_clean

a chosen behavioural metric measured on the clean and corrupted runs.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Browse the mathematical compendium →