← Mathematical compendium

Published equation contexts

z=TopKk ⁣(Wenc(a−bdec))z = \mathrm{TopK}_k\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}})\right)

Why this formula appears here

The ℓ1\ell_1 penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally, z=TopKk ⁣(Wenc(a−bdec))z = \mathrm{TopK}_k\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}})\right). keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size…

Read the full article-specific guide →

Read the representative guide

WencW_{\mathrm{enc}}

Symbol W_enc

WeW_enc is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
bdecb_{\mathrm{dec}}

Symbol b_dec

bdb_dec is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

z=TopKk ⁣(Wenc(a−bdec)),z = \mathrm{TopK}_k\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}})\right),

Equation 13 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The ℓ1\ell_1 penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally, z=TopKk ⁣(Wenc(a−bdec))z = \mathrm{TopK}_k\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}})\right). keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size…

Meanings in this article

  • aa: the writing.
Equation guide → · Article →