← Back to article

Equation 23 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

C\mathcal C

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

CC

Symbol C

sufficient to restore the behaviour.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether C\mathcal C is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether C\mathcal C is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled…
Read the full surrounding passage
Heimersheim and Nanda’s methodological account of this technique, written from direct practical experience running it, is the clearest documented source on how easily the result changes under choices that are rarely reported in full. Patching in one direction — copying clean activations into a corrupted run, “denoising” — tests whether C\mathcal C is sufficient to restore the behaviour; patching in the other direction — copying corrupted activations into a clean run, “noising” — tests whether C\mathcal C is necessary to sustain it, and the two are not mirror images of the same fact about the network. The choice of corruption itself, whether zero-ablation, Gaussian noise, or a resampled alternative prompt, changes which components appear causally load-bearing, and the choice of behavioural metric — a logit difference, a probability, a KL divergence — can shift the ranked importance of components even when every other choice is held fixed. Their central methodological warning is that a patching result is a claim about what happens when this specific substitution is made, under this specific metric and this specific corruption; it is not automatically a claim about how the component functions in general, and readers who are given a localisation result without the protocol that produced it have not been given a fully specified result [ 11 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

See this formula across 4 published contexts →

Browse the mathematical compendium →