Equation 12 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
Symbol k
What this part means
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol k→Article meaning
The passage around this formula
The penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally,
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders ↗
- [10] Interpreting Attention Layer Outputs with Sparse Autoencoders ↗
These citations provide research context; check each source for the exact claim it supports.