← Mathematical compendium

Published equation contexts

ℓ1\ell_1

Why this formula appears here

and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft ℓ1\ell_1 penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues introduced this variant, established scaling laws relating autoencoder size and sparsity to reconstruction error, and — this is the operative fact for a cost accounting — trained a sixteen-million-latent autoencoder on GPT-4’s activations over forty billion tokens [ 1 ] .

Read the full article-specific guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (8)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

ℓ1\ell_1

Equation 6 · AI Research

What Interpretability Actually Costs to Do at Scale

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft ℓ1\ell_1 penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues introduced this variant, established scaling laws relating autoencoder size and sparsity to reconstruction error, and — this is the operative fact for a cost accounting — trained a sixteen-million-latent autoencoder on GPT-4’s activations over forty billion tokens [ 1 ] .

Equation guide → · Article →
ℓ1\ell_1

Equation 10 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Equation guide → · Article →
ℓ1\ell_1

Equation 11 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Equation guide → · Article →
ℓ1\ell_1

Equation 14 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Equation guide → · Article →
ℓ1\ell_1

Equation 15 · AI Research

Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Earlier versions of this objective penalised the code’s ℓ1\ell_1 norm as a differentiable stand-in for sparsity, but an ℓ1\ell_1 penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff θ\theta trained through a straight-through gradient estimator, which lets the objective penalise the true count of active features, ∥\lVert f(x) ∥0\rVert_0 , directly rather than through the ℓ1\ell_1 proxy, and reported state-of-the-art reconstruction fidelity at matched sparsity against both the earlier…

Equation guide → · Article →
ℓ1\ell_1

Equation 10 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The reconstruction term asks the dictionary to explain the activation; the ℓ1\ell_1 penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions substantially more interpretable than neurons or principal components, and — the operational payoff — that the recovered directions support finer-grained causal attribution of specific behaviours than the alternatives available at the time [ 6 ] . Anthropic’s dictionary-learning demonstration on a one-layer model, published the same season, is the paper most responsible for making this the default first move in a new interpretability project rather than one…

Equation guide → · Article →
ℓ1\ell_1

Equation 11 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The ℓ1\ell_1 penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally,

Equation guide → · Article →
ℓ1\ell_1

Equation 17 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size grows [ 8 ] . Rajamanoharan and colleagues took a different route to the same problem, keeping a continuous encoder but replacing the fixed zero threshold with a learned per-feature threshold θi\theta_i — a feature only activates once its pre-activation clears θi\theta_i — and report state-of-the-art reconstruction fidelity at matched sparsity on Gemma 2 activations against both the ℓ1\ell_1 and top- k alternatives [ 9 ] .

Equation guide → · Article →