AI Research
The Hardest Unsolved Problems in Mechanistic Interpretability
Sparse autoencoders were meant to finish what superposition started. By the field's own papers, five problems remain open: unstable feature geometry, no ground truth, arbitrary dictionaries, and no agreed bar for a finished explanation.