Equation 3 · What Actually Happens Inside a Very Long Claude Context Window
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the model’s hidden dimension. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where d is the model’s hidden dimension. Doubling the context does not double the work of relating everything in it to everything else; it roughly quadruples that specific cost. Anthropic’s guidance frames the practical consequence in the language of a finite resource: models have what it describes as a limited attention budget, and every additional token spends a sliver of it, so a model’s capacity to represent all the pairwise relationships in a very long input “gets stretched thin” well before the window is technically full [ 4 ] . The guidance names two further contributing factors distinct from the raw attention-cost argument: models see comparatively few long sequences during training…
Read the full surrounding passage
where d is the model’s hidden dimension. Doubling the context does not double the work of relating everything in it to everything else; it roughly quadruples that specific cost. Anthropic’s guidance frames the practical consequence in the language of a finite resource: models have what it describes as a limited attention budget, and every additional token spends a sliver of it, so a model’s capacity to represent all the pairwise relationships in a very long input “gets stretched thin” well before the window is technically full [ 4 ] . The guidance names two further contributing factors distinct from the raw attention-cost argument: models see comparatively few long sequences during training relative to short ones, leaving fewer specialised parameters for context-wide dependencies, and the position-encoding schemes that let a model handle sequences longer than anything seen in training necessarily reduce the model’s precision about exactly where in a long sequence a given token sits [ 4 ] . Put together, Anthropic characterises context rot as “a performance gradient rather than a hard cliff” — a real, current-generation Claude behaviour, described in its own engineering team’s words, not a defect specific to a competitor’s system, and not a problem a strictly bigger window automatically fixes [ 4 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Actually Happens Inside a Very Long Claude Context Window