← All parts of this equation

Equation 4 · Part 11 · What a Circuit Explains: The State and Limits of Mechanistic Interpretability

subscript

L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd.\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d .
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

Now the careful part. Consider what the training objective actually asks for. Writing x for an activation vector, f(x) for the sparse code and x^\hat{x} for the reconstruction, the objective has the form L(x)=∥x−x^(x)∥22+λ∥f(x)∥1,x^(x)=Wd f(x)+bd\mathcal{L}(x) = \lVert x - \hat{x}(x) \rVert_2^2 + \lambda \lVert f(x) \rVert_1, \qquad \hat{x}(x) = W_d\, f(x) + b_d . Every term refers to the activation vector. No term refers to what the model does with that activation afterwards. The objective rewards a code that reconstructs the activation sparsely; it is indifferent to whether the dictionary elements correspond to anything the network’s downstream layers treat as a unit. Low reconstruction error at high sparsity is therefore evidence that the activation distribution is sparsely decomposable in the trained basis. It is not evidence…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.