← Back to article

Equation 14 · OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows

What does this equation mean?

sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp})

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

sadjs_{\mathrm{adj}}

Symbol s_adj

sas_adj is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

sreporteds_{\mathrm{reported}}

Symbol s_reported

the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

The size of that drop is close to what a simple deflation model predicts. If sreporteds_{\mathrm{reported}} is the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes, then sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp}). gives 0.7880 ×\times (1 - 0.1971) ≈\approx 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets…
Read the full surrounding passage
The size of that drop is close to what a simple deflation model predicts. If sreporteds_{\mathrm{reported}} is the score under the original oracle and fp\mathrm{fp} is the false-positive rate the adversarial suite exposes, then sadj≈sreported×(1−fp)s_{\mathrm{adj}} \approx s_{\mathrm{reported}} \times (1-\mathrm{fp}). gives 0.7880 ×\times (1 - 0.1971) ≈\approx 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets stricter without the model changing at all, the rank order of a leaderboard can change with it. Neither vendor’s headline SWE-bench Verified number has been re-run through this specific adversarial suite in public as of this writing, which means neither number’s resistance to the effect SWE-ABS documents is currently known.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows

See this formula across 1 published context →

Browse the mathematical compendium →