Equation 4 · Part 11 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and for the reconstruction, the objective has the form . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the dictionary elements correspond to anything the network’s downstream layers treat as a unit. Low reconstruction error at high sparsity is therefore evidence that the activation distribution is sparsely decomposable in the trained basis. It is not evidence…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the article section
- [7] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet ↗
- [10] Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers ↗
- [11] Are Sparse Autoencoders Useful? A Case Study in Sparse Probing ↗
These citations provide research context; check each source for the exact claim it supports.