Equation 5 · Part 3 · What Interpretability Actually Costs to Do at Scale
Symbol hatx
What this part means
hatx is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
hatx is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol hatx→Article meaning
The passage around this formula
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…
Learn the underlying idea
A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.
Open the illustrated functions: inputs become outputs guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.