← All parts of this equation

Equation 4 · Part 2 · How Mechanistic Interpretability Research Is Actually Done

Symbol y

p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b},\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\},
yy

What this part means

y is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

y is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

…separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation aℓa_\ell read out at layer ℓ\ell and a binary property y , p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b}\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\}. with θ\theta fit by ordinary gradient descent to minimise cross-entropy against labelled examples. Alain and Bengio introduced this move under the name “probes” and made an observation that still organises how the technique is used: linear separability of a target property…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.