Symbol hat p_θ
hat p_θ is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer …
hat p_θ is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →y is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →ll is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →σ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →op is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →b is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →θ is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →w is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 4 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation read out at layer …
Equation guide → · Article →