Symbol M
M is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →Published equation contexts
The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then…
M is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →G is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →L appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →the accuracy of that same model’s underlying text-only language backbone.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 4 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The first is measurement-time ablation : take an already-trained system, rerun the exact same benchmark items with one channel removed at inference, and compare. This is what the original VQA paper did with its question-only baseline [ 1 ] , and it is what a 2024 audit of vision-language benchmarks did formally, by defining two paired metrics from three separately measured accuracies. Let be a model’s accuracy on a benchmark with the image present, its accuracy on the same items with the image withheld, and the accuracy of that same model’s underlying text-only language backbone, evaluated on its own before any multimodal training touched it. The audit’s two metrics are then…