| Model | Most defensible read | Evidence confidence |
|---|---|---|
| Claude Opus 4.7 | Best public case for coding, software agents and multi-step work. Anthropic reports 0.715 on an internal research-agent benchmark, while Vals AI places it first on SWE-bench at 82.00% . | Medium-high |
| GPT-5.5 | Very strong general reasoning profile. O-Mega reports 92.4% on MMLU, 93.6% on GPQA Diamond, 85.0% on ARC-AGI-2 and 95.0% on ARC-AGI-1 . | Medium |
| DeepSeek V4 / V4 Pro | Promising for coding and technical experimentation, but the evidence mixes V4, V4 Pro and V4 Pro High rather than one clean model line . | |
| Kimi K2.6 | Worth watching, but not yet covered deeply enough for a full comparison. LLM Stats lists it at 0.91 on GPQA and WhatLLM includes it in a top 10 Quality Index list . | Low |
| Benchmark or metric | Claude Opus 4.7 | GPT-5.5 | DeepSeek V4 / V4 Pro | Kimi K2.6 | Practical read |
|---|---|---|---|---|---|
| SWE-bench | 82.00% on Vals AI, updated April 24, 2026 | No comparable figure in the available sources | 81% claimed by NxCode for DeepSeek V4 | No comparable figure | The cleanest public signal favors Claude. |
| SWE-bench Verified | 87.6% from Vellum; 83.5% ± 1.7 from LMCouncil | No comparable figure | Listed in a Hugging Face community evaluation for DeepSeek-V4-Pro, but no visible figure in the recovered summary | No comparable figure | Strong for Claude, but figures vary by source and setup. |
| SWE-bench Pro | 64.3% from Vellum | No comparable figure | Listed in the Hugging Face community evaluation, but no visible figure in the recovered summary | No comparable figure | More relevant to long-horizon software agents than simpler coding tests. |
| GPQA Diamond | 94.2% in O-Mega, Vellum and TNW | ||||
| MMLU | No comparable figure in the available sources | 92.4% in O-Mega | MMLU-Pro appears in community evaluation, but no visible comparable figure | No comparable figure | Low deciding power because MMLU is saturated among top models . |
| ARC-AGI | No comparable figure | ARC-AGI-2: 85.0%; ARC-AGI-1: 95.0% in O-Mega | No comparable figure | No comparable figure | Reinforces GPT-5.5 as a reasoning contender, with source caution. |
| Research-agent / multi-step work | 0.715 in Anthropic internal benchmark | No comparable figure | BenchLM reports 83.8/100 in Agentic for DeepSeek V4 Pro High | No comparable figure | Useful directional signals, but not equivalent metrics. |
| Long context / Needle-in-a-Haystack | Anthropic says Opus 4.7 had the most consistent long-context performance among models it tested | No comparable figure | NxCode reports 97% at 1M tokens, while framing it as a claim that needs independent validation | No comparable figure | DeepSeek has an interesting claim, not a settled conclusion. |
| LiveCodeBench / Codeforces | No comparable figure | No comparable figure | Redreamality reports LiveCodeBench 93.5 and Codeforces 3206 for DeepSeek V4 | No comparable figure | Positive for pure coding, but not enough to settle agentic software work. |
SWE-bench is one of the more useful coding signals because it tests whether models can resolve real-world software engineering tasks, and Vals AI describes its SWE-bench page as measuring production software engineering tasks . But SWE-bench, SWE-bench Verified and SWE-bench Pro should not be treated as the same exam. SWE-bench Pro is described in its paper as a substantially more challenging benchmark for long-horizon software engineering tasks .
GPQA Diamond is valuable for graduate-level scientific reasoning, but it is no longer a clean separator at the frontier. TNW notes that models such as Opus 4.7, GPT-5.4 Pro and Gemini 3.1 Pro are so close on GPQA Diamond that differences fall within measurement noise . MMLU needs even more caution: Nanonets says top models in 2026 are already above 88%, making the benchmark too saturated to separate leaders reliably .
Source quality also matters. An official lab post, an independent leaderboard, an aggregator and a community discussion do not carry the same weight. BenchLM, for example, says its Claude Opus 4.7 profile is excluded from the public leaderboard because it does not yet have enough non-generated public benchmark coverage to rank safely . That is a useful reminder: even strong models can have uneven public evidence.
Claude Opus 4.7 is the best-supported model in this comparison. Anthropic says Opus 4.7 tied for the top overall score across six modules in its internal research-agent benchmark at 0.715 and delivered the most consistent long-context performance among the models it tested . Because that is an internal benchmark, it should not be read as an independent leaderboard. It does, however, show where Anthropic is positioning the model: multi-step, tool-heavy work.
The cleaner outside signal is software engineering. Vals AI ranks Claude Opus 4.7 first on SWE-bench with 82.00% on a page updated April 24, 2026 . Vellum reports 87.6% on SWE-bench Verified and 64.3% on SWE-bench Pro . LMCouncil lists 83.5% ± 1.7 for Claude Opus 4.7 on SWE-bench Verified .
The right conclusion is not to pick one number and ignore the others. The careful read is that Claude appears at or near the top across multiple software-engineering views, while the exact percentage depends on benchmark variant, date, configuration and source .
On scientific reasoning, Claude Opus 4.7 is also strong: O-Mega, Vellum and TNW all show 94.2% on GPQA Diamond . But GPQA is too compressed among top models to make Claude the overall winner by itself . Claude’s more defensible edge is applied coding and agentic work.
GPT-5.5 looks like the strongest challenger on broad reasoning. O-Mega reports 92.4% on MMLU, 93.6% on GPQA Diamond, 85.0% on ARC-AGI-2 and 95.0% on ARC-AGI-1 . Vellum also lists GPT-5.5 at 93.6% on GPQA Diamond, just below Claude Opus 4.7 in that table . BenchLM places GPT-5.5 in the top tier, with an 89/100 provisional score and rank 2 of 16 on its verified leaderboard .
The caution is traceability. In the available material for this comparison, GPT-5.5 appears in articles, leaderboards and aggregator pages, but not with a full official OpenAI benchmark card comparable to Anthropic’s Opus 4.7 release material. Appwrite describes GPT-5.5 as shipped on April 23, 2026, and Vals lists openai/gpt-5.5 with a release date of April 23, 2026 and a Vals Index of 67.76% ± 1.79 . Those are useful signals, but they do not replace a first-party benchmark card.
For a decision memo, GPT-5.5 should be presented as a first-tier reasoning model, especially because of its GPQA and ARC-AGI numbers . It should not be declared the overall winner if the standard is consistent public evidence across all four models.
DeepSeek is the hardest model family to summarize cleanly because the sources move between DeepSeek V4, DeepSeek V4 Pro and DeepSeek V4 Pro High. A score for one variant should not be silently transferred to another .
Hugging Face shows a community discussion for DeepSeek-V4-Pro that adds evaluation results across GPQA, GSM8K, HLE, MMLU-Pro, SWE-bench Pro, SWE-bench Verified and Terminal-Bench 2.0 . BenchLM reports DeepSeek V4 Pro High at 83.8/100 in Agentic, 88.8/100 in Coding and 72.1/100 in Knowledge . NxCode claims DeepSeek V4 reaches 81% on SWE-bench and 97% on Needle-in-a-Haystack at 1M tokens, while also framing the long-context figure as something that needs to hold up under independent testing .
Redreamality adds another positive coding signal, reporting LiveCodeBench 93.5 and Codeforces 3206 for DeepSeek V4 . But the same source says closed frontier models still lead on long-horizon agentic work such as SWE-bench Pro and Terminal-Bench 2.0 .
The practical read: DeepSeek V4/V4 Pro deserves an internal bake-off, especially if open-weight experimentation or technical control is part of the brief. But based on the sources here, it does not yet have the same public evidence quality as Claude Opus 4.7 for SWE-bench and agentic software work .
Kimi K2.6 should not be ignored, but it should not be treated as if it has the same benchmark coverage as Claude, GPT-5.5 or DeepSeek. LLM Stats lists Kimi K2.6 at 0.91 on GPQA, and WhatLLM includes Kimi K2.6 in its top 10 models by Quality Index . Those are useful signals of benchmark activity, not enough for a broad model-to-model verdict.
It is also important not to swap in Kimi K2.5 data by accident. Simon Willison’s February 2026 SWE-bench update includes Kimi K2.5, but that is a different model version and should not be used as Kimi K2.6 evidence . For a rigorous comparison, Kimi K2.6 belongs in the pending validation column.
| Use case | Best current recommendation | Confidence | Why |
|---|---|---|---|
| Resolving real software issues and coding agents | Claude Opus 4.7 | Medium-high | Vals AI ranks it first on SWE-bench at 82.00%, and Vellum reports strong SWE-bench Verified and SWE-bench Pro results . |
| Multi-step research-agent work | Claude Opus 4.7 | Medium | Anthropic reports 0.715 on its internal research-agent benchmark and the strongest long-context consistency among models it tested . |
| Scientific reasoning in GPQA-style tasks | Claude Opus 4.7 or GPT-5.5 | Medium | Claude appears at 94.2% and GPT-5.5 at 93.6%, but the benchmark is tightly clustered among leading models . |
| Broad general reasoning | GPT-5.5 | Medium-low | Its MMLU, GPQA and ARC-AGI numbers are strong, but the available evidence is mainly from O-Mega, Vellum and BenchLM . |
| Open-weight or technical experimentation | DeepSeek V4 / V4 Pro | Medium-low | The model family has community and aggregator signals, but variant mixing and independent validation remain issues . |
| Full quantitative ranking against the other three | Do not treat Kimi K2.6 as fully comparable yet | Low | Available signals include GPQA 0.91 and a WhatLLM top 10 placement, but the coverage is too narrow . |
If you need a defensible 2026 benchmark narrative, put Claude Opus 4.7 first for coding and agentic software work. It combines an official Anthropic signal, first place on Vals AI’s SWE-bench page and strong third-party results on SWE-bench Verified and SWE-bench Pro .
Put GPT-5.5 next as the strongest broad reasoning rival. Its O-Mega and Vellum numbers are excellent, especially on GPQA and ARC-AGI, but the available evidence is less official and less uniform than Claude’s .
Treat DeepSeek V4/V4 Pro as a serious candidate for internal testing, not as a proven overall leader. The numbers are promising, but the model variants and source types need careful labeling . Treat Kimi K2.6 as insufficiently validated for a full comparison until more comparable, multi-benchmark evidence is available .
| Medium-low |
| 93.6% in O-Mega and Vellum |
| Mentioned in community suites, but no comparable visible number in the recovered summary |
| 0.91 in LLM Stats |
| Claude and GPT-5.5 are too close for GPQA alone to decide the winner. |