← Back to article

Equation 2 · What Actually Happens Inside a Very Long Claude Context Window

What does this equation mean?

Cattn(n)=O(n2⋅d),C_{\mathrm{attn}}(n) = O(n^2 \cdot d),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsO(n^2 × d)
Result or conditionC_attn(n)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

CattnC_{\mathrm{attn}}

Symbol C_attn

CaC_attn is part of the quantity the equation computes from the expression on the right.

Understand this part →

nn

Symbol n

the number of tokens.

Understand this part →

OO

Symbol O

O is one factor in the product that computes the quantity on the left.

Understand this part →

n2n^2

Symbol n^2

the square of n; the number of tokens.

Understand this part →

dd

Symbol d

the model’s hidden dimension.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Anthropic’s own engineering guidance, published in September 2025 as advice for developers building long-running agents, gives the clearest available first-party account of why context rot happens architecturally rather than treating it as an unexplained empirical curiosity [ 4 ] . The explanation rests on the transformer’s core mechanism: every token attends to every other token in the context through self-attention, so the number of pairwise relationships the model must represent grows with the square of the sequence length. For a context of n tokens, the compute spent by self-attention within a single layer scales as Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d). where d is the model’s hidden dimension. Doubling…
Read the full surrounding passage
Anthropic’s own engineering guidance, published in September 2025 as advice for developers building long-running agents, gives the clearest available first-party account of why context rot happens architecturally rather than treating it as an unexplained empirical curiosity [ 4 ] . The explanation rests on the transformer’s core mechanism: every token attends to every other token in the context through self-attention, so the number of pairwise relationships the model must represent grows with the square of the sequence length. For a context of n tokens, the compute spent by self-attention within a single layer scales as Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d). where d is the model’s hidden dimension. Doubling the context does not double the work of relating everything in it to everything else; it roughly quadruples that specific cost. Anthropic’s guidance frames the practical consequence in the language of a finite resource: models have what it describes as a limited attention budget, and every additional token spends a sliver of it, so a model’s capacity to represent all the pairwise relationships in a very long input “gets stretched thin” well before the window is technically full [ 4 ] . The guidance names two further contributing factors distinct from the raw attention-cost argument: models see comparatively few long sequences during training relative to short ones, leaving fewer specialised parameters for context-wide dependencies, and the position-encoding schemes that let a model handle sequences longer than anything seen in training necessarily reduce the model’s precision about exactly where in a long sequence a given token sits [ 4 ] . Put together, Anthropic characterises context rot as “a performance gradient rather than a hard cliff” — a real, current-generation Claude behaviour, described in its own engineering team’s words, not a defect specific to a competitor’s system, and not a problem a strictly bigger window automatically fixes [ 4 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What Actually Happens Inside a Very Long Claude Context Window

See this formula across 1 published context →

Browse the mathematical compendium →