← Back to article

Equation 1 · Cognitive Bias Was the Label; Attention Sink Is the Suspect

What does this equation mean?

SPEM=JS(P^ ∥ R)\mathrm{SPEM} = \mathrm{JS}(\hat P \,\|\, R)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsJS(hat P | R)
Result or conditionSPEM
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

P^\hat P

Symbol hat P

the predicted label distribution.

Understand this part →

RR

Symbol R

the reference.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Before any of that, credit belongs where the paper actually earned it, and the measurement is the place it earned it most cleanly. For each classification query, the paper’s own protocol produces two prompts with identical content and different label orders, records which label the model selects under each, and repeats this across three thousand randomly sampled queries per dataset — “designed to eliminate any potential bias in dataset sampling,” in the paper’s own description [ 1 ] . A primacy effect is declared when the first third of a label list accounts for more than forty percent of the model’s selections, aggregated across those three thousand trials; recency and middle effects use…
Read the full surrounding passage
Before any of that, credit belongs where the paper actually earned it, and the measurement is the place it earned it most cleanly. For each classification query, the paper’s own protocol produces two prompts with identical content and different label orders, records which label the model selects under each, and repeats this across three thousand randomly sampled queries per dataset — “designed to eliminate any potential bias in dataset sampling,” in the paper’s own description [ 1 ] . A primacy effect is declared when the first third of a label list accounts for more than forty percent of the model’s selections, aggregated across those three thousand trials; recency and middle effects use the same forty-percent rule against the last and middle thirds respectively, and their absence is scored as no effect at all. The intensity of whatever effect is found is then quantified separately, as the Jensen–Shannon divergence between the observed selection distribution and a uniform reference distribution — the paper’s own SPEM statistic, SPEM\mathrm{SPEM} = JS(P^ ∥ R)\mathrm{JS}(\hat P \,\|\, R) , where P^\hat P is the predicted label distribution and R is the reference. None of that construction is casual, and none of it is unique to one model family: thirteen checkpoints across three architecturally distinct lineages, eight datasets ranging from a seventy-seven-intent banking corpus to five-article news summaries, and a headline count — seventy-three of a hundred and four tested configurations showing primacy — that is exactly what falls out of multiplying thirteen models by eight datasets, not a rounded or approximate figure.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect

See this formula across 1 published context →

Browse the mathematical compendium →