← All parts of this equation

Equation 4 · Part 6 · How Mechanistic Interpretability Research Is Actually Done

Symbol b

p^θ(y∣aℓ)=σ ⁣(w⊤aℓ+b),θ={w,b},\hat p_\theta(y \mid a_\ell) = \sigma\!\left(w^{\top} a_\ell + b\right), \qquad \theta = \{w, b\},
bb

What this part means

b is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

b is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

The first concrete operation in almost any interpretability project is the least glamorous: run the model forward on a batch of inputs, and at some chosen point in the computation — a residual-stream position, an attention head’s output, a particular MLP layer — copy the activation tensor out before it is overwritten by the next step of the forward pass. This is extraction, and it produces nothing on its own beyond a large table of vectors. What turns it into evidence is probing: fitting a small, separately trained classifier to predict some property of interest directly from those vectors, while the model’s own weights stay frozen. Formally, for an activation aℓa_\ell read out at layer ℓ\ell…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.