← All parts of this equation

Equation 13 · Part 2 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Symbol x

∥f(x)∥0\lVert f(x) \rVert_0
xx

What this part means

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.