Equation 14 · Part 1 · How Mechanistic Interpretability Research Is Actually Done
Symbol k
What this part means
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol k→Article meaning
The passage around this formula
keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size grows [ 8 ] . Rajamanoharan and colleagues took a different route to the same problem, keeping a continuous encoder but replacing the fixed zero threshold with a learned per-feature threshold — a feature only activates once its pre-activation clears — and report state-of-the-art reconstruction fidelity at matched sparsity on Gemma 2 activations against both the and top- k alternatives [ 9 ] .
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [8] Scaling and evaluating sparse autoencoders ↗
- [9] Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders ↗
These citations provide research context; check each source for the exact claim it supports.