As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, problClaude’s reported coding-benchmark gains are substantial, but benchmark methodology is essential context for interpreting the scores.
AI 提示
Create a landscape editorial hero image for this Studio Global article: As of September 2026, how do Anthropic’s Claude models perform across the major SWE-bench coding benchmarks—including the definitions, probl. Article summary: As of September 2026, Claude appears to lead the reported SWE-bench results, but the evidence supports a narrower conclusion than “Anthropic is unambiguously best at software engineering.” SWE-bench Verified is close to . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
openai.com
Claude 的编程基准成绩曲线确实很陡:升级版 Claude 3.5 Sonnet 在 2025 年初于 SWE-bench Verified 达到 49.0%;数月后,Claude 4 系列报告为约 72.5%—72.7%;到 2026 年,部分排行榜已给出接近基准上限的成绩。23249