Equation 24 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…
Read the full surrounding passage
Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled alternative prompt, changes which components appear causally load-bearing, and the choice of behavioural metric — a logit difference, a probability, a KL divergence — can shift the ranked importance of components even when every other choice is held fixed. Their central methodological warning is that a patching result is a claim about what happens when this specific substitution is made, under this specific metric and this specific corruption; it is not automatically a claim about how the component functions in general, and readers who are given a localisation result without the protocol that produced it have not been given a fully specified result [ 11 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.