← All parts of this equation

Equation 20 · Part 2 · What Interpretability Actually Costs to Do at Scale

Symbol n^0.65

Θ(n0.65)\Theta(n^{0.65})
n0.65n^{0.65}

What this part means

n0n^0.65 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

n0n^0.65 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as Θ\Theta(n0.65n^{0.65}) tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.