← All parts of this equation

Equation 4 · Part 3 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability

Symbol hatx

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd.\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d .
x^\hat{x}

What this part means

hatx is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

hatx is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and x^\hat{x} for the reconstruction, the objective has the form L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.