← Back to article

Equation 3 · What Actually Happens Inside a Very Long Claude Context Window

What does this equation mean?

dd

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the model’s hidden dimension. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

dd

Symbol d

the model’s hidden dimension.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where d is the model’s hidden dimension. Doubling the context does not double the work of relating everything in it to everything else; it roughly quadruples that specific cost. Anthropic’s guidance frames the practical consequence in the language of a finite resource: models have what it describes as a limited attention budget, and every additional token spends a sliver of it, so a model’s capacity to represent all the pairwise relationships in a very long input “gets stretched thin” well before the window is technically full [ 4 ] . The guidance names two further contributing factors distinct from the raw attention-cost argument: models see comparatively few long sequences during training…
Read the full surrounding passage
where d is the model’s hidden dimension. Doubling the context does not double the work of relating everything in it to everything else; it roughly quadruples that specific cost. Anthropic’s guidance frames the practical consequence in the language of a finite resource: models have what it describes as a limited attention budget, and every additional token spends a sliver of it, so a model’s capacity to represent all the pairwise relationships in a very long input “gets stretched thin” well before the window is technically full [ 4 ] . The guidance names two further contributing factors distinct from the raw attention-cost argument: models see comparatively few long sequences during training relative to short ones, leaving fewer specialised parameters for context-wide dependencies, and the position-encoding schemes that let a model handle sequences longer than anything seen in training necessarily reduce the model’s precision about exactly where in a long sequence a given token sits [ 4 ] . Put together, Anthropic characterises context rot as “a performance gradient rather than a hard cliff” — a real, current-generation Claude behaviour, described in its own engineering team’s words, not a defect specific to a competitor’s system, and not a problem a strictly bigger window automatically fixes [ 4 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What Actually Happens Inside a Very Long Claude Context Window

Browse the mathematical compendium →