← Back to article

Equation 4 · What Interpretability Actually Costs to Do at Scale

What does this equation mean?

kk

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the because published topk configurations keep. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

kk

Symbol k

the because published topk configurations keep.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A sparse autoencoder (SAE) is trained to reconstruct a model’s internal activation vectors through a sparse bottleneck: an encoder maps an activation of dimension d into a much wider space of n candidate “features,” a sparsity constraint keeps only k of those features active per token, and a decoder reconstructs the original activation from just those k . The training objective, in its standard form, is

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to What Interpretability Actually Costs to Do at Scale

Browse the mathematical compendium →