← Mathematical compendium

Published equation contexts

x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0

Why this formula appears here

The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0. Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the…

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0.\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0.

Equation 9 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : x^=Wdf(x)+bd,f(x)=JumpReLUθ(Wex+be),L(x)=∥x−x^∥22+λ∥f(x)∥0\hat x = W_d f(x) + b_d, \qquad f(x) = \mathrm{JumpReLU}_\theta\big(W_e x + b_e\big), \qquad \mathcal L(x) = \lVert x - \hat x \rVert_2^2 + \lambda \lVert f(x) \rVert_0. Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the…

Equation guide → · Article →