← Mathematical compendium

Published equation contexts

Leff(τ)=max⁡{ n≤W : S(n)≥τ⋅S0 }L_{\mathrm{eff}}(\tau) = \max\left\{\, n \le W \ :\ S(n) \ge \tau \cdot S_0 \,\right\}

Why this formula appears here

Put the documentation and the independent benchmarks side by side and a pattern emerges that is more informative than any single score. Formalise the distinction the whole comparison rests on: let W be the context window a vendor documents for a given model, a fixed engineering figure set by attention implementation, position encoding, and what the serving stack supports. Let S(n) be some benchmark’s measured accuracy at input length n ≤\le W , and S0S_0 the same benchmark’s accuracy at a short reference length. For a chosen retention threshold τ\tau , define the effective context length as Leff(τ)=max⁡{ n≤W : S(n)≥τ⋅S0 }L_{\mathrm{eff}}(\tau) = \max\left\{\, n \le W \ :\ S(n) \ge \tau \cdot S_0 \,\right\}. W is a documented constant, published on day one of a model’s release, identical no…

Read the full article-specific guide →

Read the representative guide

LeffL_{\mathrm{eff}}

Symbol L_eff

a measured variable — specific to a task, a threshold, a benchmark design, and a date — and every result surveyed above shows it running well below W well before W is reached.

Read this term in its guide →
τ\tau

Symbol τ

τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Leff(τ)=max⁡{ n≤W : S(n)≥τ⋅S0 }L_{\mathrm{eff}}(\tau) = \max\left\{\, n \le W \ :\ S(n) \ge \tau \cdot S_0 \,\right\}

Equation 6 · Model Evaluation

Context Window Size Versus What a Frontier Model Can Actually Recall From It

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Put the documentation and the independent benchmarks side by side and a pattern emerges that is more informative than any single score. Formalise the distinction the whole comparison rests on: let W be the context window a vendor documents for a given model, a fixed engineering figure set by attention implementation, position encoding, and what the serving stack supports. Let S(n) be some benchmark’s measured accuracy at input length n ≤\le W , and S0S_0 the same benchmark’s accuracy at a short reference length. For a chosen retention threshold τ\tau , define the effective context length as Leff(τ)=max⁡{ n≤W : S(n)≥τ⋅S0 }L_{\mathrm{eff}}(\tau) = \max\left\{\, n \le W \ :\ S(n) \ge \tau \cdot S_0 \,\right\}. W is a documented constant, published on day one of a model’s release, identical no…

Meanings in this article

  • LeffL_{\mathrm{eff}}: a measured variable — specific to a task, a threshold, a benchmark design, and a date — and every result surveyed above shows it running well below W well before W is reached.
  • S0S_0: the same benchmark’s accuracy at a short reference length.
Equation guide → · Article →