← All parts of this equation

Equation 9 · Part 13 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

subtraction

x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0.\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0.
subtraction

What this part means

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Its job in the formula

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

The passage around this formula

The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0. Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.