Equation 9 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hat x
hat x is part of the quantity the equation computes from the expression on the right.
Symbol W_d
is one of the signed contributions combined to compute the quantity on the left.
Symbol f
f is one of the signed contributions combined to compute the quantity on the left.
Symbol x
x is part of the quantity the equation computes from the expression on the right.
Symbol b_d
is one of the signed contributions combined to compute the quantity on the left.
Symbol θ
θ is one of the signed contributions combined to compute the quantity on the left.
Symbol W_e
is one of the signed contributions combined to compute the quantity on the left.
Symbol b_e
is one of the signed contributions combined to compute the quantity on the left.
Symbol L
L is one of the signed contributions combined to compute the quantity on the left.
Symbol λ
λ is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : . Earlier versions of this objective penalised the code’s norm as a differentiable stand-in for sparsity, but an penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff trained through a straight-through gradient estimator, which lets the…
Read the full surrounding passage
The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : . Earlier versions of this objective penalised the code’s norm as a differentiable stand-in for sparsity, but an penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, f(x) , directly rather than through the proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier formulation and a competing gated variant [ 7 ] . That the field kept revising the sparsity penalty is itself evidence of how much the objective’s exact shape affects what gets recovered.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.