← Back to article

Equation 16 · Cognitive Bias Was the Label; Attention Sink Is the Suspect

What does this equation mean?

−7.5794-7.5794

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

It should not be read alone. The companion regressions for the middle and recency effects, reported directly beside it in the same appendix table set, show a different and more telling pattern. The middle-effect regression reports a Model T5 coefficient of -6.2258 carrying a standard error of 3.19 ×\times 10^{4} , and a Model Llama2 coefficient of -7.5794 carrying a standard error of 1.07 ×\times 10^{4} — standard errors between one and five thousand times the size of the coefficients they attach to. That ratio is not statistical noise; it is the textbook signature of quasi-complete separation, the condition in which a logistic regression’s maximum-likelihood estimate fails to converge because…
Read the full surrounding passage
It should not be read alone. The companion regressions for the middle and recency effects, reported directly beside it in the same appendix table set, show a different and more telling pattern. The middle-effect regression reports a Model T5 coefficient of -6.2258 carrying a standard error of 3.19 ×\times 10^{4} , and a Model Llama2 coefficient of -7.5794 carrying a standard error of 1.07 ×\times 10^{4} — standard errors between one and five thousand times the size of the coefficients they attach to. That ratio is not statistical noise; it is the textbook signature of quasi-complete separation, the condition in which a logistic regression’s maximum-likelihood estimate fails to converge because some category has too few outcomes on one side of the split for the model to be identified at all. The authors’ own conclusion, stated plainly in the same section, is admirably honest about the ambiguity: “This could be either due to our limited sample size or show that these factors are not predictive of the SPE in LLMs” [ 1 ] . That sentence is doing more work than it appears to. A regression that cannot converge cannot rule an architectural account out; at best it fails to rule it in, on a design — thirteen models from three architecturally distinct families, each contributing a single data point per dataset, with architecture family and parameter count strongly collinear by construction — that was never going to have the statistical power to separate “this doesn’t matter” from “this couldn’t be tested.” The question the appendix asks is the right one. Its own numbers show the tool it reached for could not answer it.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect

See this formula across 1 published context →

Browse the mathematical compendium →