Equation 2 · A History of How We Learned to Evaluate AI Agents
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol s_t
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
an unweighted mean of nine task scores . Averaging like this treats all nine component tasks as equally important and, implicitly, as equally hard and equally reliably measured — an assumption nothing in the construction actually guarantees. It is a real simplifying assumption, not a neutral summary statistic, and it is exactly the assumption GLUE’s own successor abandoned.
Sources cited in the article section
- [3] SQuAD: 100,000+ Questions for Machine Comprehension of Text ↗
- [4] GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding ↗
- [5] SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems ↗
- [9] Dynabench: Rethinking Benchmarking in NLP ↗
These citations give research context. Read each source to check which claims it supports.