Equation 3 · Ten Failure Modes That Define Multimodal AI Systems
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol s_text
the actual multimodal increment — what the removed channel contributed beyond what text alone already bought.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The finding motivates a simple, checkable decomposition rather than a new metric: for any multimodal benchmark, let be the score with all modalities present and be the same benchmark scored with the non-text modality removed. The quantity . is the actual multimodal increment — what the removed channel contributed beyond what text alone already bought. A large with a small is not evidence of multimodal competence; it is evidence that the benchmark, not the model, is doing something unimodal. Reporting alongside costs one extra evaluation run and turns an unfalsifiable headline number into a…
Read the full surrounding passage
The finding motivates a simple, checkable decomposition rather than a new metric: for any multimodal benchmark, let be the score with all modalities present and be the same benchmark scored with the non-text modality removed. The quantity . is the actual multimodal increment — what the removed channel contributed beyond what text alone already bought. A large with a small is not evidence of multimodal competence; it is evidence that the benchmark, not the model, is doing something unimodal. Reporting alongside costs one extra evaluation run and turns an unfalsifiable headline number into a checkable one.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to Ten Failure Modes That Define Multimodal AI Systems