Equation 5 · Part 10 · What Interpretability Actually Costs to Do at Scale
subtraction
subtraction
What this part means
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
Its job in the formula
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
Full expression→subtraction→Article meaning
The passage around this formula
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…
Learn the underlying idea
Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.
Open the illustrated addition and subtraction in an equation guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.