← All parts of this equation

Equation 25 · Part 5 · Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared

Symbol x_i^+

vℓ=1∣P∣∑i∈Phℓ(xi+)  −  1∣N∣∑i∈Nhℓ(xi−),hℓ←hℓ+α vℓ,v_\ell = \frac{1}{|P|}\sum_{i \in P} h_\ell(x_i^{+}) \;-\; \frac{1}{|N|}\sum_{i \in N} h_\ell(x_i^{-}), \qquad h_\ell \leftarrow h_\ell + \alpha\, v_\ell,
xi+x_i^{+}

What this part means

xi+x_i^+ is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

xi+x_i^+ is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

The construction is a mean-difference direction, added back into the residual stream with a tunable strength at inference time: vℓ=1∣P∣∑i∈Phℓ(xi+)  −  1∣N∣∑i∈Nhℓ(xi−),hℓ←hℓ+α vℓv_\ell = \frac{1}{|P|}\sum_{i \in P} h_\ell(x_i^{+}) \;-\; \frac{1}{|N|}\sum_{i \in N} h_\ell(x_i^{-}), \qquad h_\ell \leftarrow h_\ell + \alpha\, v_\ell. where P and N index matched positive and negative examples of the target behaviour, hℓh_\ell is the residual-stream activation at layer ℓ\ell , and α\alpha is a coefficient the operator sets by hand. Nothing in this construction requires understanding why the network represents the concept along this direction, only that adding it produces the intended behavioural shift.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.