Equation 4 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
Symbol hat p_θ
What this part means
hat p_θ is part of the quantity the equation computes from the expression on the right.
Its job in the formula
hat p_θ is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol hat p_θ→Article meaning
The passage around this formula
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer …
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.