Symbol L
L is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…
L is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →hatx is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →λ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →f is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…
Equation guide → · Article →Equation 4 · Interpretability
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and for the reconstruction, the objective has the form . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the dictionary elements correspond to anything the network’s downstream layers treat as a unit. Low reconstruction error at high sparsity is therefore evidence that the activation distribution is sparsely decomposable in the trained basis. It is not evidence…