← Back to article

Equation 12 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

What does this equation mean?

θ\theta

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

θ\theta

Symbol θ

θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…
Read the full surrounding passage
Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier ℓ1\ell_1 formulation and a competing gated variant [ 7 ] . That the field kept revising the sparsity penalty is itself evidence of how much the objective’s exact shape affects what gets recovered.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Browse the mathematical compendium →