← Mathematical compendium

Published equation contexts

C\mathcal C

Why this formula appears here

where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (4)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

C\mathcal C

Equation 21 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Meanings in this article

  • CC: copied from one run into the other.
Equation guide → · Article →
C\mathcal C

Equation 22 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

where mcleanm_{\text{clean}} and mcorruptm_{\text{corrupt}} are a chosen behavioural metric measured on the clean and corrupted runs, and mpatch(C)m_{\text{patch}}(\mathcal C) is that same metric after the activations of component set C\mathcal C are copied from one run into the other. A score near one means patching C\mathcal C alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.

Meanings in this article

  • CC: copied from one run into the other.
Equation guide → · Article →
C\mathcal C

Equation 23 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether C\mathcal C is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether C\mathcal C is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…

Meanings in this article

  • CC: sufficient to restore the behaviour.
Equation guide → · Article →
C\mathcal C

Equation 24 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether C\mathcal C is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether C\mathcal C is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…

Meanings in this article

  • CC: sufficient to restore the behaviour.
Equation guide → · Article →