Equation 5 · What Interpretability Actually Costs to Do at Scale
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L
L is part of the quantity the equation computes from the expression on the right.
Symbol x
x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol hatx
hatx is one of the signed contributions combined to compute the quantity on the left.
Symbol λ
λ is one of the signed contributions combined to compute the quantity on the left.
Symbol f
f is one of the signed contributions combined to compute the quantity on the left.
Symbol W_d
is one of the signed contributions combined to compute the quantity on the left.
Symbol b_d
is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues…
Read the full surrounding passage
A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is . and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues introduced this variant, established scaling laws relating autoencoder size and sparsity to reconstruction error, and — this is the operative fact for a cost accounting — trained a sixteen-million-latent autoencoder on GPT-4’s activations over forty billion tokens [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Interpretability Actually Costs to Do at Scale