Equation 9 · Part 5 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Symbol b_d
What this part means
is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol b_d→Article meaning
The passage around this formula
The formal objective has evolved since the earliest versions, and the direction of that evolution is itself informative. A dictionary decoder reconstructs the activation x from a sparse code f(x) : . Earlier versions of this objective penalised the code’s norm as a differentiable stand-in for sparsity, but an penalty also shrinks the magnitude of every active feature, distorting reconstruction in a way that has nothing to do with how many features are active. Rajamanoharan and colleagues introduced JumpReLU, a thresholded activation function with a learned per-feature cutoff trained through a straight-through gradient estimator, which lets the…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.