← All parts of this equation

Equation 4 · Part 1 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability

Symbol L

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd.\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d .
L\mathcal{L}

What this part means

L is part of the quantity the equation computes from the expression on the right.

Its job in the formula

L is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and x^\hat{x} for the reconstruction, the objective has the form L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the dictionary elements correspond to anything the network’s downstream layers treat as a unit. Low reconstruction error at high sparsity is therefore evidence that the activation distribution is sparsely decomposable in the trained basis. It is not evidence…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.