Equation 7 · What Interpretability Actually Costs to Do at Scale
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the because published topk configurations keep. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
and the specific variant OpenAI’s interpretability team used to push this to frontier scale, the TopK autoencoder, replaces the soft penalty with an explicit constraint: exactly k latents fire, chosen by magnitude, and the rest are hard-zeroed. Gao and colleagues introduced this variant, established scaling laws relating autoencoder size and sparsity to reconstruction error, and — this is the operative fact for a cost accounting — trained a sixteen-million-latent autoencoder on GPT-4’s activations over forty billion tokens [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Interpretability Actually Costs to Do at Scale