Published equation contexts
Why this formula appears here
where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Read the representative guide
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (4)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 21 · AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Meanings in this article
Equation guide → · Article →Equation 22 · AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
where and are a chosen behavioural metric measured on the clean and corrupted runs, and is that same metric after the activations of component set are copied from one run into the other. A score near one means patching alone restores nearly all of the clean behaviour; a score near zero means it restores almost none.
Meanings in this article
Equation guide → · Article →Equation 23 · AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…
Meanings in this article
Equation guide → · Article →Equation 24 · AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…