Equation 4 · Part 7 · The Benchmarks That Don't Need the Image
=
=
What this part means
The expressions on both sides represent the same quantity under the stated assumptions.
Its job in the formula
The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.
Full expression→=→Article meaning
The passage around this formula
The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then…
Learn the underlying idea
An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.
Open the illustrated equality: what the equals sign claims guide →
Sources cited in the surrounding passage
- [1] VQA: Visual Question Answering ↗
- [10] Are We on the Right Way for Evaluating Large Vision-Language Models? ↗
These citations provide research context; check each source for the exact claim it supports.