Claude 在 SWE bench Verified 的公開成績,從 2025 年初升級版 Claude 3.5 Sonnet 的 49.0%,進展至榜單所列 Claude Opus 5 的 97.0%。 SWE bench Verified 僅有 500 道公開、以 Python 為主的真實 GitHub 問題,接近飽和後,分數已較難區分前沿系統;SWE bench Pro 覆蓋更廣,但結果也高度受代理框架與統計方式影響。
Claude 在 SWE bench Verified 的公開成績,從 2025 年初升級版 Claude 3.5 Sonnet 的 49.0%,進展至榜單所列 Claude Opus 5 的 97.0%。
SWE bench Verified 僅有 500 道公開、以 Python 為主的真實 GitHub 問題,接近飽和後,分數已較難區分前沿系統;SWE bench Pro 覆蓋更廣,但結果也高度受代理框架與統計方式影響。
Terminal Bench 4.0 的一項引用比較中,Claude Fable 5.1 為 55.8%,GPT 5.6 Sol 為 37.3%;這支持該組態在此項評測領先,卻不足以構成所有 AI 供應商的全面排名。
As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, problClaude’s reported coding-benchmark gains are substantial, but benchmark methodology is essential context for interpreting the scores.
AI 提示詞
Create a landscape editorial hero image for this Studio Global article: As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, probl. Article summary: As of September 2026, Claude appears to lead the reported SWE-bench results, but the evidence supports a narrower conclusion than “Anthropic is unambiguously best at software engineering.” SWE-bench Verified is close to . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
openai.com
Claude 的程式基準成績確實出現大幅躍升:2025 年初,升級版 Claude 3.5 Sonnet 在 SWE-bench Verified 達到 49.0%;數月後,Claude 4 系列來到約 72.5% 至 72.7%;後續榜單則出現接近該基準上限的結果。23249
到了 2025 年 5 月,Anthropic 報告 Claude Opus 4 為 72.5%、Claude Sonnet 4 為 72.7%;兩項數字皆是在未使用延伸思考(extended thinking)的條件下取得。24
至 2026 年,榜單情勢已明顯不同。Vals AI 列出 Claude Opus 5 為 97.0%,並指出有 7 個受評模型達到 95% 以上,DeepSeek V4 Pro 0813 也有 96.4%。9 因此,把 Opus 5 說成「約 96%」可視為對高九成分數的概略說法,卻不是唯一權威數字:模型版本、代理框架、預算與評測設定,都會改變結果。
Claude 在 SWE bench Verified 的公開成績,從 2025 年初升級版 Claude 3.5 Sonnet 的 49.0%,進展至榜單所列 Claude Opus 5 的 97.0%。
最值得優先驗證的重點是什麼?
Claude 在 SWE bench Verified 的公開成績,從 2025 年初升級版 Claude 3.5 Sonnet 的 49.0%,進展至榜單所列 Claude Opus 5 的 97.0%。 SWE bench Verified 僅有 500 道公開、以 Python 為主的真實 GitHub 問題,接近飽和後,分數已較難區分前沿系統;SWE bench Pro 覆蓋更廣,但結果也高度受代理框架與統計方式影響。
接下來在實務上該怎麼做?
Terminal Bench 4.0 的一項引用比較中,Claude Fable 5.1 為 55.8%,GPT 5.6 Sol 為 37.3%;這支持該組態在此項評測領先,卻不足以構成所有 AI 供應商的全面排名。