← Mathematical compendium

Published equation contexts

∥f(x)∥0\lVert f(x) \rVert_0

Why this formula appears here

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Read the full article-specific guide →

Read the representative guide

ff

Symbol f

f is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

∥f(x)∥0\lVert f(x) \rVert_0

Equation 13 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Equation guide → · Article →