← Mathematical compendium

Published equation contexts

CSAE≈2dnTC_{\mathrm{SAE}} \approx 2dnT

Why this formula appears here

Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as Θ\Theta(n0.65n^{0.65}) tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…

Read the full article-specific guide →

Read the representative guide

CSAEC_{\mathrm{SAE}}

Symbol C_SAE

CSC_SAE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
dd

Symbol d

d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
nn

Symbol n

n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
TT

Symbol T

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (2)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

CSAE≈2dnTC_{\mathrm{SAE}} \approx 2dnT

Equation 18 · AI Research

What Interpretability Actually Costs to Do at Scale

This equation gives an approximation: it relates the quantities while allowing an approximation.

Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n ≫\gg k and the encoding term dominates almost entirely: CSAEC_{\mathrm{SAE}} ≈\approx 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as Θ\Theta(n0.65n^{0.65}) tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…

Equation guide → · Article →
CSAE≈2dnTC_{\mathrm{SAE}} \approx 2dnT

Equation 26 · AI Research

What Interpretability Actually Costs to Do at Scale

This equation gives an approximation: it relates the quantities while allowing an approximation.

The compute ledger has a plausible, if unproven, path downward. CSAEC_{\mathrm{SAE}} ≈\approx 2dnT is a cost that infrastructure and algorithmic improvements — sparser encoders, better initialization that reaches convergence at lower n , transcoders that replace rather than merely observe a component — can attack directly, the same way serving-side engineering rather than raw parameter growth has driven most within-generation price reduction for inference elsewhere in this field. TopK autoencoders were themselves exactly this kind of improvement over the softer, less efficient sparsity penalties they replaced [ 1 ] . There is no comparable engineering lever visible yet for the labor ledger.…

Equation guide → · Article →