Equation 6 · Part 3 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
That gap between “predictable from” and “used by” is the method’s central and openly documented weakness. Hewitt and Liang showed that a sufficiently expressive probe can achieve high accuracy predicting properties from representations that plausibly do not encode them in any meaningful sense, because the probe itself has the capacity to memorise idiosyncratic patterns in the training data. Their fix was the control task: construct a version of the labelling scheme that associates each input type with an output at random, so it can only be learned by the probe’s own memorisation capacity, never by any real linguistic signal in the representation. A well-behaved probe should then show high…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the surrounding passage
- [3] Designing and Interpreting Probes with Control Tasks ↗
- [4] Probing Classifiers: Promises, Shortcomings, and Advances ↗
These citations provide research context; check each source for the exact claim it supports.