← All parts of this equation

Equation 20 · Part 3 · What Interpretability Actually Costs to Do at Scale

superscript

Θ(n0.65)\Theta(n^{0.65})
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as Θ\Theta(n0.65n^{0.65}) tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.