← Back to article

Equation 1 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

y^(x)=σ(w⊤hℓ(x)+b),\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsσbig(w^top h_ell(x) + bbig)
Result or conditionhat y(x)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

y^\hat y

Symbol hat y

hat y is part of the quantity the equation computes from the expression on the right.

Understand this part →

xx

Symbol x

x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

σ\sigma

Symbol σ

σ is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

w⊤w^\top

Symbol w^top

wtw^top is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

hℓh_\ell

Symbol h_ell

the frozen activation the network produces at layer ℓ\ell for input x , and w.

Understand this part →

bb

Symbol b

the probe’s own parameters, trained on labelled examples the network never saw during its own training.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Formally, a probe is a function fitted independently of the network under study: y^(x)=σ(w⊤hℓ(x)+b)\hat y(x) = \sigma\big(w^\top h_\ell(x) + b\big). where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

See this formula across 1 published context →

Browse the mathematical compendium →