Equation 25 · Part 5 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Symbol x_i^+
What this part means
is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol x_i^+→Article meaning
The passage around this formula
The construction is a mean-difference direction, added back into the residual stream with a tunable strength at inference time: . where P and N index matched positive and negative examples of the target behaviour, is the residual-stream activation at layer , and is a coefficient the operator sets by hand. Nothing in this construction requires understanding why the network represents the concept along this direction, only that adding it produces the intended behavioural shift.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [12] Representation Engineering: A Top-Down Approach to AI Transparency ↗
- [13] Steering Language Models With Activation Engineering ↗
- [14] Steering Llama 2 via Contrastive Activation Addition ↗
- [15] Analyzing the Generalization and Reliability of Steering Vectors ↗
These citations provide research context; check each source for the exact claim it supports.