Equation 4 · Part 12 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
superscript
superscript
What this part means
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
Its job in the formula
A raised mark can be a power or an index. Its position and the surrounding notation determine which.
Full expression→superscript→Article meaning
The passage around this formula
Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and for the reconstruction, the objective has the form . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the dictionary elements correspond to anything the network’s downstream layers treat as a unit. Low reconstruction error at high sparsity is therefore evidence that the activation distribution is sparsely decomposable in the trained basis. It is not evidence…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
Sources cited in the article section
- [7] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet ↗
- [10] Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers ↗
- [11] Are Sparse Autoencoders Useful? A Case Study in Sparse Probing ↗
These citations provide research context; check each source for the exact claim it supports.