← All parts of this equation

Equation 4 · Part 3 · The Benchmarks That Don't Need the Image

Symbol S_v

MG=Sv−Swv,ML=max⁡(0, Swv−St)MG = S_v - S_{wv}, \qquad ML = \max(0,\ S_{wv} - S_t)
SvS_v

What this part means

SvS_v appears in the objective or constraint used by the optimization on the right.

Its job in the formula

SvS_v appears in the objective or constraint used by the optimization on the right.

The passage around this formula

…the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.