← All parts of this equation

Equation 3 · Part 1 · The Benchmarks That Don't Need the Image

Symbol S_t

StS_t
StS_t

What this part means

the accuracy of that same model’s underlying text-only language backbone.

Its job in the formula

StS_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it.

The passage around this formula

The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let SvS_v be a model’s accuracy on a benchmark with the image present, SwvS_{wv} its accuracy on the same items with the image withheld, and StS_t the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.