Equation 10 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
The reconstruction term asks the dictionary to explain the activation; the penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions substantially more interpretable than neurons or principal components, and — the operational payoff — that the recovered directions support finer-grained causal attribution of specific behaviours than the alternatives available at the time [ 6 ] . Anthropic’s dictionary-learning demonstration on a one-layer model, published the same season, is the paper most responsible for making this the default first move in a new interpretability project rather than one…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the surrounding passage
- [6] Sparse Autoencoders Find Highly Interpretable Features in Language Models ↗
- [7] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
These citations provide research context; check each source for the exact claim it supports.