Equation 1 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hat y
hat y is part of the quantity the equation computes from the expression on the right.
Symbol x
x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol σ
σ is one of the signed contributions combined to compute the quantity on the left.
Symbol w^top
op is one of the signed contributions combined to compute the quantity on the left.
Symbol h_ell
the frozen activation the network produces at layer for input x , and w.
Symbol b
the probe’s own parameters, trained on labelled examples the network never saw during its own training.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Formally, a probe is a function fitted independently of the network under study: . where is the frozen activation the network produces at layer for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.
Sources cited in the article section
- [1] Understanding Intermediate Layers Using Linear Classifier Probes ↗
- [2] What You Can Cram into a Single Vector: Probing Sentence Embeddings for Linguistic Properties ↗
- [3] Designing and Interpreting Probes with Control Tasks ↗
- [4] Probing Classifiers: Promises, Shortcomings, and Advances ↗
These citations give research context. Read each source to check which claims it supports.