← All parts of this equation

Equation 1 · Part 7 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

=

y^(x)=σ(w⊤hℓ(x)+b),\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big),
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

Formally, a probe is a function fitted independently of the network under study: y^(x)=σ(w⊤hℓ(x)+b)\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big). where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.