← Mathematical compendium

Published equation contexts

selectivity=acctask−acccontrol\text{selectivity} = \mathrm{acc}_{\text{task}} - \mathrm{acc}_{\text{control}}

Why this formula appears here

That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high…

Read the full article-specific guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

selectivity=acctask−acccontrol.\text{selectivity} = \mathrm{acc}_{\text{task}} - \mathrm{acc}_{\text{control}}.

Equation 6 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high…

Equation guide → · Article →