Equation 16 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol x
x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The decisive limitation is not about engineering the penalty correctly; it is about what the objective, however penalised, actually certifies. Every term in that loss function refers only to the activation vector x and its reconstruction. No term refers to what the network’s own downstream layers do with that activation. Low reconstruction error at high sparsity is evidence that the activation distribution happens to be sparsely decomposable in the basis the autoencoder found. It is not evidence that this is the basis the network itself treats as its working units. Chanin and colleagues gave this a sharp, concrete demonstration under the name feature absorption: when the true underlying…
Read the full surrounding passage
The decisive limitation is not about engineering the penalty correctly; it is about what the objective, however penalised, actually certifies. Every term in that loss function refers only to the activation vector x and its reconstruction. No term refers to what the network’s own downstream layers do with that activation. Low reconstruction error at high sparsity is evidence that the activation distribution happens to be sparsely decomposable in the basis the autoencoder found. It is not evidence that this is the basis the network itself treats as its working units. Chanin and colleagues gave this a sharp, concrete demonstration under the name feature absorption: when the true underlying features form a hierarchy — a general property and a more specific one that implies it — the sparsity objective can cause the general feature to stop firing on cases where it logically should, because those cases have been “absorbed” into the more specific child feature instead. The result is a dictionary that looks clean and monosemantic on inspection while silently failing to fire where a human reading of the concept would expect it to, precisely because sparsity, not fidelity to the network’s own computation, is what the objective rewards [ 8 ] . A sparse autoencoder answers “what is a low-dimensional, sparse basis that reconstructs this activation,” not “what are this network’s own computational units,” and the two questions can have different correct answers.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.