Symbol C_SAE
AE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
Work out what that architecture actually spends compute on per token, because the two halves of it behave differently. The encoding step needs a score for every one of the n candidate latents before it can select the top k , so it is an unavoidably dense matrix multiply: roughly 2dn floating-point operations. The decoding step only touches the k latents that survived, so it is sparse: roughly 2dk operations. Summed and multiplied across T training tokens, a first-order compute model for training the dictionary is . Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n k and the encoding term dominates almost…
AE is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 14 · AI Research
This equation gives an approximation: it relates the quantities while allowing an approximation.
Work out what that architecture actually spends compute on per token, because the two halves of it behave differently. The encoding step needs a score for every one of the n candidate latents before it can select the top k , so it is an unavoidably dense matrix multiply: roughly 2dn floating-point operations. The decoding step only touches the k latents that survived, so it is sparse: roughly 2dk operations. Summed and multiplied across T training tokens, a first-order compute model for training the dictionary is . Because published TopK configurations keep k in the tens to low hundreds while n runs into the millions, n k and the encoding term dominates almost…