← All parts of this equation

Equation 1 · Part 4 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Symbol w^top

y^(x)=σ(w⊤hℓ(x)+b),\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big),
w⊤w^\top

What this part means

wtw^top is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

wtw^top is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

Formally, a probe is a function fitted independently of the network under study: y^(x)=σ(w⊤hℓ(x)+b)\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big). where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.