← All parts of this equation

Equation 1 · Part 8 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

addition

y^(x)=σ(w⊤hℓ(x)+b),\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big),
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

Formally, a probe is a function fitted independently of the network under study: y^(x)=σ(w⊤hℓ(x)+b)\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big). where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.