← All parts of this equation

Equation 14 · Part 5 · What Interpretability Actually Costs to Do at Scale

Symbol T

CSAE  ≈  2 d (n+k) T.C_{\mathrm{SAE}} \;\approx\; 2\,d\,(n + k)\,T .
TT

What this part means

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

…k , so it is an unavoidably dense matrix multiply: roughly 2dn floating-point operations. The decoding step only touches the k latents that survived, so it is sparse: roughly 2dk operations. Summed and multiplied across T training tokens, a first-order compute model for training the dictionary is CSAE  ≈  2 d (n+k) TC_{\mathrm{SAE}} \;\approx\; 2\,d\,(n + k)\,T . Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.