Equation 4 · Part 3 · The Benchmarks That Don't Need the Image
Symbol S_v
What this part means
appears in the objective or constraint used by the optimization on the right.
Its job in the formula
appears in the objective or constraint used by the optimization on the right.
Full expression→Symbol S_v→Article meaning
The passage around this formula
…the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [1] VQA: Visual Question Answering ↗
- [10] Are We on the Right Way for Evaluating Large Vision-Language Models? ↗
These citations provide research context; check each source for the exact claim it supports.