Equation 25 · Part 19 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Starting index or lower bound: i in N
What this part means
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Its job in the formula
i in N appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Starting index or lower bound: i in N→Article meaning
The passage around this formula
The construction is a mean-difference direction, added back into the residual stream with a tunable strength at inference time: . where P and N index matched positive and negative examples of the target behaviour, is the residual-stream activation at layer , and is a coefficient the operator sets by hand. Nothing in this construction requires understanding why the network represents the concept along this direction, only that adding it produces the intended behavioural shift.
Learn the underlying idea
Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.
Open the illustrated sums and products: repeat an operation over an index guide →
Sources cited in the article section
- [12] Representation Engineering: A Top-Down Approach to AI Transparency ↗
- [13] Steering Language Models With Activation Engineering ↗
- [14] Steering Llama 2 via Contrastive Activation Addition ↗
- [15] Analyzing the Generalization and Reliability of Steering Vectors ↗
These citations provide research context; check each source for the exact claim it supports.