← All parts of this equation

Equation 9 · Part 1 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Symbol hat x

x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0.\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0.
x^\hat x

What this part means

hat x is part of the quantity the equation computes from the expression on the right.

Its job in the formula

hat x is part of the quantity the equation computes from the expression on the right.

The passage around this formula

The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0. Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.