← Back to article

Equation 3 · Ten Failure Modes That Define Multimodal AI Systems

What does this equation mean?

Δ=smulti−stext\Delta = s_{\text{multi}} - s_{\text{text}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationss_multi - s_text
Result or conditionΔ
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Δ\Delta

Symbol Δ

the small.

Understand this part →

smultis_{\text{multi}}

Symbol s_multi

the large.

Understand this part →

stexts_{\text{text}}

Symbol s_text

the actual multimodal increment — what the removed channel contributed beyond what text alone already bought.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The finding motivates a simple, checkable decomposition rather than a new metric: for any multimodal benchmark, let smultis_{\text{multi}} be the score with all modalities present and stexts_{\text{text}} be the same benchmark scored with the non-text modality removed. The quantity Δ=smulti−stext\Delta = s_{\text{multi}} - s_{\text{text}}. is the actual multimodal increment — what the removed channel contributed beyond what text alone already bought. A large smultis_{\text{multi}} with a small Δ\Delta is not evidence of multimodal competence; it is evidence that the benchmark, not the model, is doing something unimodal. Reporting Δ\Delta alongside smultis_{\text{multi}} costs one extra evaluation run and turns an unfalsifiable headline number into a…
Read the full surrounding passage
The finding motivates a simple, checkable decomposition rather than a new metric: for any multimodal benchmark, let smultis_{\text{multi}} be the score with all modalities present and stexts_{\text{text}} be the same benchmark scored with the non-text modality removed. The quantity Δ=smulti−stext\Delta = s_{\text{multi}} - s_{\text{text}}. is the actual multimodal increment — what the removed channel contributed beyond what text alone already bought. A large smultis_{\text{multi}} with a small Δ\Delta is not evidence of multimodal competence; it is evidence that the benchmark, not the model, is doing something unimodal. Reporting Δ\Delta alongside smultis_{\text{multi}} costs one extra evaluation run and turns an unfalsifiable headline number into a checkable one.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Ten Failure Modes That Define Multimodal AI Systems

See this formula across 1 published context →

Browse the mathematical compendium →