← All parts of this equation

Equation 5 · Part 3 · What Interpretability Actually Costs to Do at Scale

Symbol hatx

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd,\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d,
x^\hat{x}

What this part means

hatx is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

hatx is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d. and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft ℓ1\ell_1 penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.