← All parts of this equation

Equation 5 · Part 9 · What Interpretability Actually Costs to Do at Scale

addition

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd,\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d,
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d. and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft ℓ1\ell_1 penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.