Equation 4 · Part 3 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
Symbol hatx
What this part means
hatx is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
hatx is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol hatx→Article meaning
The passage around this formula
Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and for the reconstruction, the objective has the form . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the…
Learn the underlying idea
A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.
Open the illustrated functions: inputs become outputs guide →
See this notation across published equations →
Sources cited in the article section
- [7] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet ↗
- [10] Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers ↗
- [11] Are Sparse Autoencoders Useful? A Case Study in Sparse Probing ↗
These citations provide research context; check each source for the exact claim it supports.