← All parts of this equation

Equation 9 · Part 3 · How Mechanistic Interpretability Research Is Actually Done

Symbol hat a

LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc).\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right).
a^\hat a

What this part means

hat a is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

hat a is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

…same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and a^\hat a for the reconstruction, LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc)\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right). The reconstruction term asks the dictionary to explain the activation; the ℓ1\ell_1 penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions substantially…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.