Symbol L_SAE
AE is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and for the reconstruction, . The reconstruction term asks the dictionary to explain the activation; the penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…
AE is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →hat a is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →λ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →z is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →ec is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →ec is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →nc is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →nc is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 9 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Sparse dictionary learning is the field’s answer, and it is a second and different act of extraction rather than a departure from the first: a sparse autoencoder is trained on the very same activation vectors the probe was reading, learning an overcomplete basis in which each vector is reconstructed as a sparse combination of dictionary elements. Writing a for the activation, z for its sparse code and for the reconstruction, . The reconstruction term asks the dictionary to explain the activation; the penalty asks it to explain it using as few active dictionary elements as possible at once. Cunningham and colleagues showed this produces directions…