Equation 9 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
Symbol L_SAE
What this part means
AE is part of the quantity the equation computes from the expression on the right.
Its job in the formula
AE is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol L_SAE→Article meaning
The passage around this formula
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and for the reconstruction, . The reconstruction term asks the dictionary to explain the activation; the penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
These citations provide research context; check each source for the exact claim it supports.