← All parts of this equation

Equation 5 · Part 7 · What Interpretability Actually Costs to Do at Scale

Symbol b_d

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd,\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d,
bdb_d

What this part means

bdb_d is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

bdb_d is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d. and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft ℓ1\ell_1 penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.