Equation 1 · Part 5 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Symbol h_ell
What this part means
the frozen activation the network produces at layer for input x , and w.
Its job in the formula
ll is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol h_ell→Article meaning
Where the article explains it
where is the frozen activation the network produces at layer for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training.
The passage around this formula
Formally, a probe is a function fitted independently of the network under study: . where is the frozen activation the network produces at layer for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [1] Understanding Intermediate Layers Using Linear Classifier Probes ↗
- [2] What You Can Cram into a Single Vector: Probing Sentence Embeddings for Linguistic Properties ↗
- [3] Designing and Interpreting Probes with Control Tasks ↗
- [4] Probing Classifiers: Promises, Shortcomings, and Advances ↗
These citations provide research context; check each source for the exact claim it supports.