Equation 6 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high…
Read the full surrounding passage
That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high accuracy on the real task and low accuracy on the matched control, and the paper defines the resulting quantity as selectivity: . A probe with low selectivity is not measuring the network; it is measuring itself [ 3 ] . Belinkov’s survey of the whole probing-classifier programme, published in Computational Linguistics , restates the deeper problem this points to: probing is correlational by construction. A property can be decodable from an activation and still play no causal role in what the network actually outputs, because the network’s downstream layers may never read that component of the representation, or may read it in combination with enough other information that the isolated correlation is misleading. The survey’s own framing is that probing classifiers have genuine promises, real and well-documented shortcomings — chiefly the presence-versus-use gap and the risk of probe overcapacity — and a set of methodological advances, control tasks among them, that partially but not fully close those gaps [ 4 ] .
Sources cited in the surrounding passage
- [3] Designing and Interpreting Probes with Control Tasks ↗
- [4] Probing Classifiers: Promises, Shortcomings, and Advances ↗
These citations give research context. Read each source to check which claims it supports.