Equation 22 · Part 1 · What Interpretability Actually Costs to Do at Scale
Symbol h_i
What this part means
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol h_i→Article meaning
The passage around this formula
A simple model makes the scaling problem legible. Let a verified circuit claim require passing m distinct checks — faithfulness under intervention, completeness against the behavior it is meant to fully explain, minimality against components that turn out not to matter — and let each check take researcher-hours, including the false starts a competent skeptic would force. Total verification cost per claim is
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [4] Auditing language models for hidden objectives ↗
- [7] In-context Learning and Induction Heads ↗
- [8] Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small ↗
- [9] Towards Automated Circuit Discovery for Mechanistic Interpretability ↗
These citations provide research context; check each source for the exact claim it supports.