OpenAI and Claude on Agentic Coding: What the Independent Evidence Actually Shows
System cards show resolve rates climbing fast on SWE-bench Verified and Terminal-Bench, but scaffold, effort, and test-oracle strength move a score as much as weights do. The one reproduction holding harness fixed separates OpenAI and Claude by under a point.