← All parts of this equation

Equation 4 · Part 7 · The Benchmarks That Don't Need the Image

=

MG=Sv−Swv,ML=max⁡(0, Swv−St)MG = S_v - S_{wv}, \qquad ML = \max(0,\ S_{wv} - S_t)
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.