← All parts of this equation

Equation 4 · Part 7 · How Mechanistic Interpretability Research Is Actually Done

Symbol θ

p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b},\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\},
θ\theta

What this part means

θ is part of the quantity the equation computes from the expression on the right.

Its job in the formula

θ is part of the quantity the equation computes from the expression on the right.

The passage around this formula

…to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation aℓa_\ell read out at layer ℓ\ell and a binary property y , p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b}\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\}. with θ\theta fit by ordinary gradient descent to minimise cross-entropy against labelled examples. Alain and Bengio introduced this move under the name “probes” and made an observation that still organises how the technique is used: linear separability of a target property increases monotonically with…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.