Equation 9 · Part 13 · How Mechanistic Interpretability Research Is Actually Done
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and for the reconstruction, . The reconstruction term asks the dictionary to explain the activation; the penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the surrounding passage
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
These citations provide research context; check each source for the exact claim it supports.