As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, problClaude’s reported coding-benchmark gains are substantial, but benchmark methodology is essential context for interpreting the scores.
AI プロンプト
Create a landscape editorial hero image for this Studio Global article: As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, probl. Article summary: As of September 2026, Claude appears to lead the reported SWE-bench results, but the evidence supports a narrower conclusion than “Anthropic is unambiguously best at software engineering.” SWE-bench Verified is close to . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
2025年5月には、AnthropicがSWE-bench Verifiedで**Claude Opus 4は72.5%、Claude Sonnet 4は72.7%**と報告した。これらは拡張思考を使わない条件での数値である。24
2026年にはリーダーボードの様相が大きく変わった。Vals AIは**Claude Opus 5を97.0%**とし、評価済みモデルのうち7モデルが95%以上、DeepSeek V4 Pro 0813も96.4%と報告している。9 したがって「Opus 5は96%程度」という表現は、高90%台の結果を指す略記としては理解できる。ただし、唯一の決定的な数値ではない。スコアはモデルの版、エージェント基盤、トークンや時間の予算、評価設定に左右される。