Equation 13 · Cognitive Bias Was the Label; Attention Sink Is the Suspect
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol p
p is part of the quantity the equation computes from the expression on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The paper’s own appendix, in a section titled “The Predictability of Serial Position Effects,” comes close to running exactly the kind of test this article is arguing for — and its result is more interesting, and more damaging to any confident reading in either direction, than either the main text or the paper’s abstract lets on. The authors fit a logistic regression predicting which type of serial position effect appears — primacy, recency, middle, or none — from four candidate features: model size in parameters, task accuracy, the rate at which a model’s predicted label changes when the list is reshuffled, and model architecture family, encoded as a set of dummy variables [ 1 ] . This is,…
Read the full surrounding passage
The paper’s own appendix, in a section titled “The Predictability of Serial Position Effects,” comes close to running exactly the kind of test this article is arguing for — and its result is more interesting, and more damaging to any confident reading in either direction, than either the main text or the paper’s abstract lets on. The authors fit a logistic regression predicting which type of serial position effect appears — primacy, recency, middle, or none — from four candidate features: model size in parameters, task accuracy, the rate at which a model’s predicted label changes when the list is reshuffled, and model architecture family, encoded as a set of dummy variables [ 1 ] . This is, in substance, an attempt to ask whether something about a model’s construction — exactly the question an attention-sink account would want answered — predicts the effect. On the MASSIVE dataset, the primacy-effect regression reports a Model Size coefficient of -0.0345 with a standard error of 0.103 and p = 0.738 : a small, statistically unremarkable coefficient, cleanly estimated, that shows no relationship. Read alone, that looks like real evidence against any scale-dependent architectural account.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect