Symbol C_SAE
AE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n k and the encoding term dominates almost entirely: 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as () tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…
AE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 18 · AI Research
This equation gives an approximation: it relates the quantities while allowing an approximation.
Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n k and the encoding term dominates almost entirely: 2dnT . That single approximation explains something the paper reports without deriving: convergence — the point at which more tokens stop buying lower reconstruction error — is reached later as n grows, empirically as () tokens for GPT-4-scale autoencoders [ 1 ] . Cost scales with the product of dictionary width and token count, and pushing width up forces token count up too if the dictionary is to be trained to convergence rather than merely trained. The paper is explicit that this collided…
Equation guide → · Article →Equation 26 · AI Research
This equation gives an approximation: it relates the quantities while allowing an approximation.
The compute ledger has a plausible, if unproven, path downward. 2dnT is a cost that infrastructure and algorithmic improvements — sparser encoders, better initialization that reaches convergence at lower n , transcoders that replace rather than merely observe a component — can attack directly, the same way serving-side engineering rather than raw parameter growth has driven most within-generation price reduction for inference elsewhere in this field. TopK autoencoders were themselves exactly this kind of improvement over the softer, less efficient sparsity penalties they replaced [ 1 ] . There is no comparable engineering lever visible yet for the labor ledger.…
Equation guide → · Article →