← Mathematical compendium

Published equation contexts

LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc)\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right)

Why this formula appears here

Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and a^\hat a for the reconstruction, LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc)\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right). The reconstruction term asks the dictionary to explain the activation; the ℓ1\ell_1 penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…

Read the full article-specific guide →

Read the representative guide

LSAE\mathcal{L}_{\mathrm{SAE}}

Symbol L_SAE

LSL_SAE is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
WdecW_{\mathrm{dec}}

Symbol W_dec

WdW_dec is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
bdecb_{\mathrm{dec}}

Symbol b_dec

bdb_dec is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
WencW_{\mathrm{enc}}

Symbol W_enc

WeW_enc is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
bencb_{\mathrm{enc}}

Symbol b_enc

beb_enc is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc).\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right).

Equation 9 · AI Research

How Mechanistic Interpretability Research Is Actually Done

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and a^\hat a for the reconstruction, LSAE(a)=∥a−a^∥22+λ∥z∥1,a^=Wdec z+bdec,z=ReLU ⁣(Wenc(a−bdec)+benc)\mathcal{L}_{\mathrm{SAE}}(a) = \lVert a - \hat a \rVert_2^2 + \lambda \lVert z \rVert_1, \qquad \hat a = W_{\mathrm{dec}}\, z + b_{\mathrm{dec}}, \qquad z = \mathrm{ReLU}\!\left(W_{\mathrm{enc}}(a - b_{\mathrm{dec}}) + b_{\mathrm{enc}}\right). The reconstruction term asks the dictionary to explain the activation; the ℓ1\ell_1 penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…

Meanings in this article

  • aa: the writing.
Equation guide → · Article →