Symbol z
z is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
The penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally, . keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size…
z is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →k is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →nc is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →ec is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 13 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The penalty has a known cost: it does not directly control how many dictionary elements fire, only how much their combined magnitude is discouraged, and it systematically shrinks the elements that do fire toward zero, biasing the reconstruction. Two later refinements address this more directly. Gao and colleagues introduced k -sparse encoding, which drops the tunable penalty in favour of a fixed sparsity budget enforced structurally, . keeping only the k largest pre-activations and zeroing the rest, and used it to train a sixteen-million-latent dictionary on GPT-4 activations over forty billion tokens, reporting metrics that improve consistently as dictionary size…