← Back to article

Equation 26 · Cognitive Bias Was the Label; Attention Sink Is the Suspect

What does this equation mean?

ϕsink\phi_{\mathrm{sink}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ϕsink\phi_{\mathrm{sink}}

Symbol phi_sink

phisi_sink is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

What survives every argument above, untouched, is the paper’s actual empirical contribution. The measurement itself — that primacy and recency sensitivity is widespread across architecturally distinct model families, that it is not confined to decoder-only transformers since the encoder-decoder T5 and Flan-T5 models showed comparable effects, that it persists into summarization tasks via BERTScore-measured focus rather than only multiple-choice selection, and that mitigation through prompting is real but unreliable — does not depend in any way on which mechanism turns out to explain it. A reader deciding whether to trust an LLM-as-judge system on an unlabeled ranking task needs to know that…
Read the full surrounding passage
What survives every argument above, untouched, is the paper’s actual empirical contribution. The measurement itself — that primacy and recency sensitivity is widespread across architecturally distinct model families, that it is not confined to decoder-only transformers since the encoder-decoder T5 and Flan-T5 models showed comparable effects, that it persists into summarization tasks via BERTScore-measured focus rather than only multiple-choice selection, and that mitigation through prompting is real but unreliable — does not depend in any way on which mechanism turns out to explain it. A reader deciding whether to trust an LLM-as-judge system on an unlabeled ranking task needs to know that position matters and how much, not why it matters, and the paper answers the “how much” question rigorously regardless of how the ϕsink\phi_{\mathrm{sink}} measurement eventually comes out. Nor does the countermodel explain everything the paper reports: the appendix’s own finding that neither model size nor architecture family significantly predicts which specific type of effect appears is, on a well-behaved sub-model like the MASSIVE primacy regression, a genuine data point that an architecture-only story has to sit with rather than explain away, even granting that the same appendix’s other two regressions were too unstable to trust either direction.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect

Browse the mathematical compendium →