Equation 14 · OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol s_adj
dj is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol s_reported
the score under the original oracle and is the false-positive rate the adversarial suite exposes.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
The size of that drop is close to what a simple deflation model predicts. If is the score under the original oracle and is the false-positive rate the adversarial suite exposes, then . gives 0.7880 (1 - 0.1971) 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets…
Read the full surrounding passage
The size of that drop is close to what a simple deflation model predicts. If is the score under the original oracle and is the false-positive rate the adversarial suite exposes, then . gives 0.7880 (1 - 0.1971) 0.633 — within a percentage point of the measured 0.622 [ 14 ] . That closeness is a coincidence of rounding as much as a proof of the model, since rejected patches are not independent of task difficulty, but the approximation is useful precisely because it names the assumption plainly: a benchmark score is a joint statement about the system under test and the strength of the judge grading it, and when the judge gets stricter without the model changing at all, the rank order of a leaderboard can change with it. Neither vendor’s headline SWE-bench Verified number has been re-run through this specific adversarial suite in public as of this writing, which means neither number’s resistance to the effect SWE-ABS documents is currently known.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows