Equation 9 · Part 2 · How Mechanistic Interpretability Research Is Actually Done
Symbol a
What this part means
the writing.
Its job in the formula
a is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Full expression→Symbol a→Article meaning
Where the article explains it
Writing a for the activation, z for its sparse code and for the reconstruction, .
The passage around this formula
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
These citations provide research context; check each source for the exact claim it supports.