← All parts of this equation

Equation 9 · Part 2 · How Mechanistic Interpretability Research Is Actually Done

Symbol a

LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc).\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right).
aa

What this part means

the writing.

Its job in the formula

a is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Where the article explains it

Writing a for the activation, z for its sparse code and a^\hat a for the reconstruction, LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc)\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right).

The passage around this formula

Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.