← All parts of this equation

Equation 13 · Part 4 · How Mechanistic Interpretability Research Is Actually Done

Symbol a

z=TopKk ⁣(Wenc(a−bdec)),z = \mathrm{TopK}_k\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}})\right),
aa

What this part means

the writing.

Its job in the formula

a is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

Writing a for the activation, z for its sparse code and a^\hat a for the reconstruction, LSAE(a)\mathcal{L}_{\mathrm{SAE}}(a) = ∥\lVert a - a^\hat a ∥22\rVert_2^2 + λ\lambda ∥\lVert z ∥1\rVert_1, \qquad a^\hat a = WdecW_{\mathrm{dec}}\, z + bdecb_{\mathrm{dec}}, \qquad z = ReLU\mathrm{ReLU}\!

The passage around this formula

The ℓ1\ell_1 penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.