Equation 11 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
The penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally,
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the article section
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders ↗
- [10] Interpreting Attention Layer Outputs with Sparse Autoencoders ↗
These citations provide research context; check each source for the exact claim it supports.