← Back to article

Equation 19 · Cognitive Bias Was the Label; Attention Sink Is the Suspect

What does this equation mean?

ϕsink=1−SPEMsigmoidSPEMsoftmax\phi_{\mathrm{sink}} = 1 - \frac{\mathrm{SPEM}_{\mathrm{sigmoid}}}{\mathrm{SPEM}_{\mathrm{softmax}}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withSPEM_sigmoid
Divide bySPEM_softmax
This relates tophi_sink
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ϕsink\phi_{\mathrm{sink}}

Symbol phi_sink

phisi_sink is part of the quantity the equation computes from the expression on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

SPEMsigmoid\mathrm{SPEM}_{\mathrm{sigmoid}}

Numerator: SPEM_sigmoid

The complete quantity above the fraction bar.

Understand this part →

SPEMsoftmax\mathrm{SPEM}_{\mathrm{softmax}}

Denominator: SPEM_softmax

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Everything above is a case for treating the architectural account as at least as well supported as the psychological framing. It is not, on its own, a decisive discriminator, and this article’s one genuine contribution is naming and specifying the measurement that would be. Call it the sink-attributable fraction, ϕsink\phi_{\mathrm{sink}} , defined as ϕsink=1−SPEMsigmoidSPEMsoftmax\phi_{\mathrm{sink}} = 1 - \frac{\mathrm{SPEM}_{\mathrm{sigmoid}}}{\mathrm{SPEM}_{\mathrm{softmax}}}. where SPEMsoftmax\mathrm{SPEM}_{\mathrm{softmax}} is Guo and Vosoughi’s own Jensen–Shannon-divergence effect magnitude, measured by their own published protocol, on an ordinary sink-forming model, and SPEMsigmoid\mathrm{SPEM}_{\mathrm{sigmoid}} is the identical measurement on an architecture- and data-matched twin trained with sigmoid attention…
Read the full surrounding passage
Everything above is a case for treating the architectural account as at least as well supported as the psychological framing. It is not, on its own, a decisive discriminator, and this article’s one genuine contribution is naming and specifying the measurement that would be. Call it the sink-attributable fraction, ϕsink\phi_{\mathrm{sink}} , defined as ϕsink=1−SPEMsigmoidSPEMsoftmax\phi_{\mathrm{sink}} = 1 - \frac{\mathrm{SPEM}_{\mathrm{sigmoid}}}{\mathrm{SPEM}_{\mathrm{softmax}}}. where SPEMsoftmax\mathrm{SPEM}_{\mathrm{softmax}} is Guo and Vosoughi’s own Jensen–Shannon-divergence effect magnitude, measured by their own published protocol, on an ordinary sink-forming model, and SPEMsigmoid\mathrm{SPEM}_{\mathrm{sigmoid}} is the identical measurement on an architecture- and data-matched twin trained with sigmoid attention instead of softmax — the exact contrast Gu and colleagues’ released training pipeline was built to produce [ 3 ] . The quantity is dimensionless, bounded in the neighborhood of zero to one for any model pair where the sink-free twin is not somehow more position-sensitive than its sink-forming counterpart, and it answers a single, sharp question: of the effect magnitude the original paper measured, how much of it disappears when the one specific architectural mechanism this article names is removed and nothing else changes?

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Cognitive Bias Was the Label; Attention Sink Is the Suspect

See this formula across 1 published context →

Browse the mathematical compendium →