← Mathematical compendium

Published equation contexts

2dk2dk

Why this formula appears here

Work out what that architecture actually spends compute on per token, because the two halves of it behave differently. The encoding step needs a score for every one of the n candidate latents before it can select the top k , so it is an unavoidably dense matrix multiply: roughly 2dn floating-point operations. The decoding step only touches the k latents that survived, so it is sparse: roughly 2dk operations. Summed and multiplied across T training tokens, a first-order compute model for training the dictionary is

Read the full article-specific guide →

Read the representative guide

dd

Symbol d

d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

2dk2dk

Equation 12 · AI Research

What Interpretability Actually Costs to Do at Scale

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Work out what that architecture actually spends compute on per token, because the two halves of it behave differently. The encoding step needs a score for every one of the n candidate latents before it can select the top k , so it is an unavoidably dense matrix multiply: roughly 2dn floating-point operations. The decoding step only touches the k latents that survived, so it is sparse: roughly 2dk operations. Summed and multiplied across T training tokens, a first-order compute model for training the dictionary is

Meanings in this article

  • kk: the because published topk configurations keep.
Equation guide → · Article →