← All parts of this equation

Equation 14 · Part 1 · What Interpretability Actually Costs to Do at Scale

Symbol C_SAE

CSAE  ≈  2 d (n+k) T.C_{\mathrm{SAE}} \;\approx\; 2\,d\,(n + k)\,T .
CSAEC_{\mathrm{SAE}}

What this part means

CSC_SAE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

CSC_SAE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

…model for training the dictionary is CSAE  ≈  2 d (n+k) TC_{\mathrm{SAE}} \;\approx\; 2\,d\,(n + k)\,T . Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as Θ\Theta(n0.65n^{0.65}) tokens for GPT-4-scale autoencoders [ 1…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.