← All parts of this equation

Equation 1 · Part 6 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Symbol b

y^(x)=σ(w⊤hℓ(x)+b),\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big),
bb

What this part means

the probe’s own parameters, trained on labelled examples the network never saw during its own training.

Its job in the formula

b is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training.

The passage around this formula

Formally, a probe is a function fitted independently of the network under study: y^(x)=σ(w⊤hℓ(x)+b)\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big). where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.