Equation 7 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
Symbol z
What this part means
z is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
z is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol z→Article meaning
The passage around this formula
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and for the reconstruction,
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders ↗
- [10] Interpreting Attention Layer Outputs with Sparse Autoencoders ↗
These citations provide research context; check each source for the exact claim it supports.