Equation 3 · How Mechanistic Interpretability Research Is Actually Done
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol y
y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer …
Read the full surrounding passage
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer and a binary property y ,
Sources cited in the article section
- [1] Understanding Intermediate Layers Using Linear Classifier Probes ↗
- [3] Probing Classifiers: Promises, Shortcomings, and Advances ↗
- [2] Designing and Interpreting Probes with Control Tasks ↗
These citations give research context. Read each source to check which claims it supports.
Return to How Mechanistic Interpretability Research Is Actually Done