← All parts of this equation

Equation 4 · Part 6 · The Benchmarks That Don't Need the Image

Symbol S_t

MG=Sv−Swv,ML=max⁡(0, Swv−St)MG = S_v - S_{wv}, \qquad ML = \max(0,\ S_{wv} - S_t)
StS_t

What this part means

the accuracy of that same model’s underlying text-only language backbone.

Its job in the formula

StS_t appears in the objective or constraint used by the optimization on the right.

Where the article explains it

Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it.

The passage around this formula

…formally, by defining two paired metrics from three separately measured accuracies. Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then MG=Sv−Swv,ML=max⁡(0, Swv−St)MG = S_v - S_{wv}, \qquad ML = \max(0,\ S_{wv} - S_t). where MG , Multimodal Gain, is how much the image actually added once everything else is held constant,…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.