Equation 3 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where is the frozen activation the network produces at layer for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.
Sources cited in the article section
- [1] Understanding Intermediate Layers Using Linear Classifier Probes ↗
- [2] What You Can Cram into a Single Vector: Probing Sentence Embeddings for Linguistic Properties ↗
- [3] Designing and Interpreting Probes with Control Tasks ↗
- [4] Probing Classifiers: Promises, Shortcomings, and Advances ↗
These citations give research context. Read each source to check which claims it supports.