← Back to article

Equation 6 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

selectivity=acctask−acccontrol.\text{selectivity} = \mathrm{acc}_{\text{task}} - \mathrm{acc}_{\text{control}}.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsacc_task - acc_control
Result or conditionselectivity
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high…
Read the full surrounding passage
That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high accuracy on the real task and low accuracy on the matched control, and the paper defines the resulting quantity as selectivity: selectivity=acctask−acccontrol\text{selectivity} = \mathrm{acc}_{\text{task}} - \mathrm{acc}_{\text{control}}. A probe with low selectivity is not measuring the network; it is measuring itself [ 3 ] . Belinkov’s survey of the whole probing-classifier programme, published in Computational Linguistics , restates the deeper problem this points to: probing is correlational by construction. A property can be decodable from an activation and still play no causal role in what the network actually outputs, because the network’s downstream layers may never read that component of the representation, or may read it in combination with enough other information that the isolated correlation is misleading. The survey’s own framing is that probing classifiers have genuine promises, real and well-documented shortcomings — chiefly the presence-versus-use gap and the risk of probe overcapacity — and a set of methodological advances, control tasks among them, that partially but not fully close those gaps [ 4 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

See this formula across 1 published context →

Browse the mathematical compendium →