Equation 4 · Part 6 · The Benchmarks That Don't Need the Image
Symbol S_t
What this part means
the accuracy of that same model’s underlying text-only language backbone.
Its job in the formula
appears in the objective or constraint used by the optimization on the right.
Full expression→Symbol S_t→Article meaning
Where the article explains it
Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it.
The passage around this formula
…formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then . where MG , Multimodal Gain, is how much the image actually added once everything else is held constant,…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [1] VQA: Visual Question Answering ↗
- [10] Are We on the Right Way for Evaluating Large Vision-Language Models? ↗
These citations provide research context; check each source for the exact claim it supports.