← All parts of this equation

Equation 2 · Part 9 · What Actually Happens Inside a Very Long Claude Context Window

superscript

Cattn(n)=O(n2⋅d),C_{\mathrm{attn}}(n) = O(n^2 \cdot d),
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Anthropic’s own engineering guidance, published in September 2025 as advice for developers building long-running agents, gives the clearest available first-party account of why context rot happens architecturally rather than treating it as an unexplained empirical curiosity [ 4 ] . The explanation rests on the transformer’s core mechanism: every token attends to every other token in the context through self-attention, so the number of pairwise relationships the model must represent grows with the square of the sequence length. For a context of n tokens, the compute spent by self-attention within a single layer scales as Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d). where d is the model’s hidden dimension. Doubling…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.