← All parts of this equation

Equation 4 · Part 3 · How Mechanistic Interpretability Research Is Actually Done

Symbol a_ell

p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b},\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\},
aℓa_\ell

What this part means

aea_ell is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

aea_ell is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

…turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation aℓa_\ell read out at layer ℓ\ell and a binary property y , p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b}\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\}. with θ\theta fit by ordinary gradient descent to minimise cross-entropy against labelled examples. Alain and Bengio introduced this move under the name “probes” and made an observation that still organises how the technique…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.