Claude progressed from a 49% SWE bench Verified result in early 2025 to reported frontier leading results by September 2026, including 97.0% for Opus 5 on one Verified leaderboard and 81.2% for Fable 5.1 on a Pro aggr... SWE bench Verified is a 500 task, public, Python focused issue resolution benchmark and is neari...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, probl. Article summary: As of September 2026, Claude appears to lead the reported SWE-bench results, but the evidence supports a narrower conclusion than “Anthropic is unambiguously best at software engineering.” SWE-bench Verified is close to . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Claude’s reported coding-benchmark trajectory is dramatic: upgraded Claude 3.5 Sonnet reached 49% on SWE-bench Verified in early 2025; Claude 4 models reached roughly 72.5%–72.7% a few months later; and later leaderboards reported scores close to the ceiling of the benchmark. 23
24
9
The right conclusion is not that Claude has solved software engineering. It is that Claude has become highly competitive—and often leading—on published coding-agent evaluations. As scores approach 100% on a finite public benchmark, the harder question becomes which evaluation best predicts performance on a team’s own repositories, tooling, and review process.
| Benchmark | What an agent does | Coverage and scoring | Why it matters |
|---|---|---|---|
| SWE-bench Verified | Receives a real GitHub issue and a repository snapshot, then produces a patch. | A human-filtered set of 500 SWE-bench instances, primarily from popular open-source Python repositories. A result is counted as resolved when the patch passes the benchmark tests in the evaluation harness. |
A widely recognized test of issue resolution in real codebases. |
| SWE-bench Pro | Solves harder, longer-horizon repository tasks under a standardized agent scaffold. | Reported as 1,865 tasks from 41 repositories spanning 123 languages; commonly scored as Pass@1. |
Intended to broaden beyond Verified’s public Python-centered tasks and reduce contamination risk. |
| Terminal-Bench 4.0 | Carries out end-to-end technical work in a terminal environment. | Includes tasks that require investigation, commands, editing, testing, and recovery rather than only a GitHub issue patch. |
Adds operational and tool-use demands that are closer to agent workflows. |
Anthropic reported that upgraded Claude 3.5 Sonnet achieved 49.0% on SWE-bench Verified, above the prior reported 45% state of the art. At that point, Anthropic noted that no model had yet passed 50% on the benchmark. 23
37
In May 2025, Anthropic reported 72.5% for Claude Opus 4 and 72.7% for Claude Sonnet 4 on SWE-bench Verified. Those figures were reported without extended thinking. 24
By 2026, the leaderboard picture had changed substantially. Vals AI reports Claude Opus 5 at 97.0%, with seven evaluated models at 95% or above and DeepSeek V4 Pro 0813 at 96.4%. 9 That makes “Opus 5 at 96%” a reasonable shorthand for a high-90s reported result, but not a uniquely authoritative number: scores depend on the model version, agent scaffold, budget, and evaluation setup.
A 97% result on a 500-task set leaves very little room to separate frontier systems. More importantly, Verified is composed of public, finite tasks from public repositories. Benchmark guidance characterizes it as high risk for contamination and saturation. 3
That does not mean the score is meaningless. It means the score is most useful as evidence that an agent can solve this class of well-specified, testable issue—and less useful as a complete forecast of novel work in a private, evolving codebase.
SWE-bench Pro is positioned as a more difficult, more contamination-resistant alternative. Reported coverage is 1,865 tasks across 41 repositories and 123 programming languages, compared with Verified’s 500 human-filtered instances from popular Python repositories. 12
16
Published September 2026 aggregate leaderboards list Claude Fable 5.1 at 81.2%, ahead of Claude Mythos 5 at 80.3% and Claude Fable 5 at 80.0%. 53 On that leaderboard, Claude holds the top three positions.
However, there is an important qualification: not all Pro numbers are directly comparable. One source distinguishes a 59.1% leading score on Scale’s standardized public set from higher vendor-aggregate figures, illustrating how much the agent harness and reporting methodology can affect the result. 13 Another source explicitly warns that the 81.2% Fable 5.1 figure is not clearly supported by a directly published model-specific SWE-bench result and may be an aggregation or version-attribution issue.
55
The practical takeaway is simple: treat 81.2% as a reported leaderboard value, not as a settled, independently replicated fact about the base model alone.
Terminal-Bench evaluates technical agent work in a command-line environment, such as building software or training a machine-learning model. It differs from SWE-bench Verified’s issue-and-patch format by testing a more operational workflow. 38
The cited comparison reports 55.8% for Claude Fable 5.1 and 37.3% for GPT-5.6 Sol, an 18.5-point gap. 44 That is meaningful evidence that the reported Claude configuration performed better on this evaluation.
It does not establish a general ranking of Anthropic, OpenAI, Google, Cognition, or AWS. A broad vendor claim would require matched models, identical agent scaffolds, comparable token and time budgets, and results across several independent task suites. The supplied evidence does not provide that comparison.
The available results support three measured claims:
They do not show that Claude will autonomously complete multi-week projects, reliably understand undocumented internal systems, make sound architecture decisions, or ship safely without review. SWE-bench tasks have verifiable test outcomes; real engineering also involves requirements discovery, trade-offs, integration, security, deployment, and collaboration.
SWE-bench Verified was valuable because it moved beyond short coding puzzles: agents work with real repositories and issue descriptions, and patches are graded with tests. 1
38 But models have rapidly advanced from around 40% to above 80% on the evaluation, according to Anthropic’s discussion of agent evals.
38
At this stage, a useful assessment stack should include:
Anthropic has described both SWE-bench Verified and Terminal-Bench as examples of agent evaluations, and it has emphasized long-running capability for Claude 4. 38
42 But the supplied evidence is not sufficient to show a formal company-wide shift away from SWE-bench toward longer-horizon evaluations. That stronger claim should remain unproven.
Claude’s reported benchmark performance makes it one of the strongest coding-agent options to evaluate as of September 2026. The clearest milestone is the move from 49% on SWE-bench Verified in early 2025 to a reported 97.0% Opus 5 leaderboard score, alongside a reported 81.2% lead on SWE-bench Pro. 23
9
53
For tool selection, do not choose from a single headline. Use these scores to shortlist agents, then run a controlled evaluation on representative tasks in your own repositories. Near-saturated public benchmarks show that models can fix many known-style issues; they cannot, by themselves, prove dependable autonomous engineering in a production organization.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Claude progressed from a 49% SWE bench Verified result in early 2025 to reported frontier leading results by September 2026, including 97.0% for Opus 5 on one Verified leaderboard and 81.2% for Fable 5.1 on a Pro aggr...
Claude progressed from a 49% SWE bench Verified result in early 2025 to reported frontier leading results by September 2026, including 97.0% for Opus 5 on one Verified leaderboard and 81.2% for Fable 5.1 on a Pro aggr... SWE bench Verified is a 500 task, public, Python focused issue resolution benchmark and is nearing saturation; SWE bench Pro is broader and designed to be more resistant to contamination, but published scores can depe...
The often cited 55.8% versus 37.3% Terminal Bench 4.0 comparison favors the reported Claude Fable 5.1 configuration over GPT 5.6 Sol, but it is one benchmark configuration—not a comprehensive vendor ranking.