← Mathematical compendium

Published equation contexts

(a+b) mod p(a+b) \bmod p

Why this formula appears here

Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b)  mod \bmod p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by…

Read the full article-specific guide →

Read the representative guide

aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
bb

Symbol b

b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
pp

Symbol p

p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

(a+b) mod p(a+b) \bmod p

Equation 23 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Nanda and colleagues trained small transformers on nothing but modular addition — predicting (a+b)  mod \bmod p for a fixed prime p — a task deliberately chosen because the textbook answer to “what algorithm does this” is fixed and checkable in advance, and studied networks that grok: memorising the training set first, with poor generalisation, and only much later transitioning sharply to a solution that generalises, often long after training loss has already bottomed out [ 15 ] . Extraction and probing came first, applied to the network’s own weights and intermediate activations, and revealed something a probe alone could not have guessed going in: the trained embeddings organised inputs by…

Equation guide → · Article →