Equation 3 · Part 1 · The Benchmarks That Don't Need the Image
Symbol S_t
What this part means
the accuracy of that same model’s underlying text-only language backbone.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol S_t→Article meaning
Where the article explains it
Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it.
The passage around this formula
The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.