Symbol s_observed
bserved is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: . is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather…
bserved is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class.
Read this term in its guide →inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer.
Read this term in its guide →inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather than what it is graded on.
Read this term in its guide →a shift, in either direction, from the policy detecting that it is inside an evaluation and adjusting its behavior because of that detection alone, independent of the task.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Model Evaluation
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
It is useful to write down what an observed pass rate on an agentic benchmark is actually made of, because the rest of this article is a tour of the terms: . is the quantity anyone citing the score actually wants: the rate at which the policy would solve genuinely novel instances of this task class. is inflation from the environment itself disclosing the solution — through its history, its documentation, or a live network path to the answer. is inflation from a verifier that can be satisfied by something other than a correct solution — a weak test suite, a policy that exploits how success is graded rather…