Equation 3 · Part 1 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
Symbol hatx
What this part means
hatx is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
hatx is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol hatx→Article meaning
The passage around this formula
Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and for the reconstruction, the objective has the form
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [7] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet ↗
- [10] Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers ↗
- [11] Are Sparse Autoencoders Useful? A Case Study in Sparse Probing ↗
These citations provide research context; check each source for the exact claim it supports.