Equation 23 · How Mechanistic Interpretability Research Is Actually Done
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol a
a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol b
b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol p
p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b) p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by…
Read the full surrounding passage
Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b) p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by their residue around a circle, one circle frequency per output logit, rather than in the sparse-lookup structure a memorising solution would use. Decomposing the network’s Fourier-domain activity — a spectral rather than a sparse-dictionary decomposition, but playing the same structural role of breaking a tangled representation into legible parts — showed a circuit computing sums and differences of trigonometric identities across those frequencies, an algorithm with a closed mathematical form that can be checked against the network’s actual weights term by term, not merely inferred from behaviour.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How Mechanistic Interpretability Research Is Actually Done