← Mathematical compendium

Published equation contexts

sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp})

Why this formula appears here

The size of that drop is close to what a simple deflation model predicts. If sreporteds_{\mathrm{reported}} is the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes, then sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp}). gives 0.7880 ×\times (1 - 0.1971) ≈\approx 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets…

Read the full article-specific guide →

Read the representative guide

sadjs_{\mathrm{adj}}

Symbol s_adj

sas_adj is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
sreporteds_{\mathrm{reported}}

Symbol s_reported

the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp})

Equation 14 · Model Evaluation

OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows

This equation gives an approximation: it relates the quantities while allowing an approximation.

The size of that drop is close to what a simple deflation model predicts. If sreporteds_{\mathrm{reported}} is the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes, then sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp}). gives 0.7880 ×\times (1 - 0.1971) ≈\approx 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets…

Meanings in this article

Equation guide → · Article →