Equation 8 · Part 1 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Symbol f
What this part means
f is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
f is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol f→Article meaning
The passage around this formula
The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) :
Learn the underlying idea
A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.
Open the illustrated functions: inputs become outputs guide →
See this notation across published equations →
Sources cited in the article section
- [5] Toy Models of Superposition ↗
- [6] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning ↗
- [7] Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders ↗
- [8] A Is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders ↗
These citations provide research context; check each source for the exact claim it supports.