Symbol a
a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b) p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by…
a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 23 · AI Research
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b) p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by…
Equation guide → · Article →