← Back to article

Equation 4 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

xx

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where hℓ(x)h_\ell(x) is the frozen activation the network produces at layer ℓ\ell for input x , and w, b are the probe’s own parameters, trained on labelled examples the network never saw during its own training. Nothing in that objective touches the network’s weights or its downstream computation. A probe reports only whether some linear function of this one activation predicts the label well.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Browse the mathematical compendium →