Equation 3 · Cognitive Bias Was the Label; Attention Sink Is the Suspect
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the reference. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Before any of that, credit belongs where the paper actually earned it, and the measurement is the place it earned it most cleanly. For each classification query, the paper’s own protocol produces two prompts with identical content and different label orders, records which label the model selects under each, and repeats this across three thousand randomly sampled queries per dataset — “designed to eliminate any potential bias in dataset sampling,” in the paper’s own description [ 1 ] . A primacy effect is declared when the first third of a label list accounts for more than forty percent of the model’s selections, aggregated across those three thousand trials; recency and middle effects use…
Read the full surrounding passage
Before any of that, credit belongs where the paper actually earned it, and the measurement is the place it earned it most cleanly. For each classification query, the paper’s own protocol produces two prompts with identical content and different label orders, records which label the model selects under each, and repeats this across three thousand randomly sampled queries per dataset — “designed to eliminate any potential bias in dataset sampling,” in the paper’s own description [ 1 ] . A primacy effect is declared when the first third of a label list accounts for more than forty percent of the model’s selections, aggregated across those three thousand trials; recency and middle effects use the same forty-percent rule against the last and middle thirds respectively, and their absence is scored as no effect at all. The intensity of whatever effect is found is then quantified separately, as the Jensen–Shannon divergence between the observed selection distribution and a uniform reference distribution — the paper’s own SPEM statistic, = , where is the predicted label distribution and R is the reference. None of that construction is casual, and none of it is unique to one model family: thirteen checkpoints across three architecturally distinct lineages, eight datasets ranging from a seventy-seven-intent banking corpus to five-article news summaries, and a headline count — seventy-three of a hundred and four tested configurations showing primacy — that is exactly what falls out of multiplying thirteen models by eight datasets, not a rounded or approximate figure.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect