OpenAI has abandoned both of its flagship coding benchmarks: it retired SWE bench Verified in February 2026 after an audit found at least 59.4% of the hardest test cases had flaws that rejected correct solutions, then...
Research answer

Create a landscape editorial hero image for this Studio Global article: Search & fact-check with cited sources for What led OpenAI to retract its recommendation of SWE-Bench Pro, and how does this retraction fit. Article summary: Here are the findings on the two-phase collapse of OpenAI's SWE-bench benchmarks and the broader pattern of AI coding benchmark erosion.. Topic tags: general, general web, user generated, education, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layo
In just seven months, OpenAI exited both of the coding benchmarks it had championed. First came the retirement of SWE-bench Verified in February 2026, a benchmark OpenAI co-created and which had become the industry standard. Then, in July 2026, OpenAI retracted its own recommendation of the intended replacement, SWE-bench Pro. This back-to-back retreat is more than a logistical problem for model evaluators — it is a signal that the way the AI industry measures coding capability is fundamentally broken.
SWE-bench Verified launched in August 2024 as a human-validated subset of the original Princeton SWE-bench, comprising 500 Python tasks drawn from real GitHub issues O. For roughly 18 months, it served as the primary yardstick for how well AI coding agents could resolve real-world software problems T.
On February 23, 2026, OpenAI's Frontier Evals team formally deprecated the benchmark TO. The reasons were stark:
OpenAI's conclusion was unambiguous: "Improvements on SWE-bench Verified no longer reflect meaningful improvements in models' real-world software development capabilities. They increasingly reflect how much the model was exposed to the benchmark at training time" O.
OpenAI explicitly recommended SWE-bench Pro — a larger benchmark built by Scale AI from private and copyleft repositories — as the replacement OT.
On July 8, 2026, OpenAI reported the results of a detailed audit of SWE-bench Pro — the very benchmark it had just promoted as more robust. The findings were devastating TG:
This forced OpenAI to retract its recommendation of SWE-bench Pro, leaving the industry without a trusted successor benchmark G.
OpenAI's two-step retreat is not an isolated mishap. It is part of a systemic crisis in how the AI field evaluates coding ability:
OpenAI's back-to-back retreat — first abandoning its own benchmark, then disowning the replacement — has left the AI coding evaluation landscape without a trusted leader. The community increasingly recognizes that high benchmark scores no longer reliably predict whether an AI coding agent can handle real-world software engineering tasks GA. New evaluation methodologies — such as task-specific, adversarial, or continuously updated benchmarks — are urgently needed but not yet mature GAA.
For now, anyone trying to evaluate an AI coding agent has no single standard to trust. The collapse of SWE-bench Verified and SWE-bench Pro is not just a story about two flawed tests. It is a story about an industry that built faster than it could measure.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI has abandoned both of its flagship coding benchmarks: it retired SWE bench Verified in February 2026 after an audit found at least 59.4% of the hardest test cases had flaws that rejected correct solutions, then...