Equation 4 · Part 8 · How Mechanistic Interpretability Research Is Actually Done
Symbol w
What this part means
w is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
w is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol w→Article meaning
The passage around this formula
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer …
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.