← Mathematical compendium

Published equation contexts

Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d)

Why this formula appears here

Anthropic’s own engineering guidance, published in September 2025 as advice for developers building long-running agents, gives the clearest available first-party account of why context rot happens architecturally rather than treating it as an unexplained empirical curiosity [ 4 ] . The explanation rests on the transformer’s core mechanism: every token attends to every other token in the context through self-attention, so the number of pairwise relationships the model must represent grows with the square of the sequence length. For a context of n tokens, the compute spent by self-attention within a single layer scales as Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d). where d is the model’s hidden dimension. Doubling…

Read the full article-specific guide →

Read the representative guide

CattnC_{\mathrm{attn}}

Symbol C_attn

CaC_attn is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Cattn(n)=O(n2⋅d),C_{\mathrm{attn}}(n) = O(n^2 \cdot d),

Equation 2 · Foundation Models

What Actually Happens Inside a Very Long Claude Context Window

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Anthropic’s own engineering guidance, published in September 2025 as advice for developers building long-running agents, gives the clearest available first-party account of why context rot happens architecturally rather than treating it as an unexplained empirical curiosity [ 4 ] . The explanation rests on the transformer’s core mechanism: every token attends to every other token in the context through self-attention, so the number of pairwise relationships the model must represent grows with the square of the sequence length. For a context of n tokens, the compute spent by self-attention within a single layer scales as Cattn(n)=O(n2⋅d)C_{\mathrm{attn}}(n) = O(n^2 \cdot d). where d is the model’s hidden dimension. Doubling…

Meanings in this article

  • nn: the number of tokens.
  • n2n^2: the square of n; the number of tokens.
  • dd: the model’s hidden dimension.
Equation guide → · Article →