Treat the scores as use case signals, not a final league table. Public results mix reasoning settings, reporting dates, provider reported numbers and third party analysis.
Research answer

Create a landscape editorial hero image for this Studio Global article: GPT-5.5・Claude Opus 4.7・DeepSeek V4・Kimi K2.6比較:ベンチマークで見る用途別の勝者. Article summary: 4モデルを完全同一条件で横比較した公開表は確認できないため、単一の勝者ではなく用途別に選ぶのが安全です。総合候補はGPT 5.5(AA Intelligence 59、GDPval AA Elo 1785)とClaude Opus 4.7(共通10ベンチマークで6勝4敗)です。[4][26][27]. Topic tags: ai, llm benchmarks, openai, anthropic, deepseek. Reference image context from search candidates: Reference image 1: visual subject "[Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison](https://www.youtube.com/watch?v=M90iB4hpenI). . [](https://www.youtube.com" source context "Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison - YouTube" Reference image 2: visual subject "[Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison](https://www.youtube.com/watch?v=M90iB4hpenI). . [](
The useful question is not Which model is best? It is Which one should run this workload? Public benchmark data for GPT-5.5, Claude Opus 4.7, DeepSeek V4 and Kimi K2.6 mixes different reasoning settings, publication dates, provider-reported scores and third-party analyses. That makes a single all-purpose ranking too neat to trust.
For DeepSeek, the clearest comparable figures here are for DeepSeek V4 Pro, Reasoning, Max Effort. Artificial Analysis lists DeepSeek V4 Pro and Kimi K2.6 side by side in its open-model table, including Intelligence score, context window, price column and output speed.
The closest head-to-head is GPT-5.5 against Claude Opus 4.7. Mashable reported the following side-by-side benchmark scores.
| Benchmark | GPT-5.5 | Claude Opus 4.7 | Lead in Mashable table |
|---|---|---|---|
| SWE-Bench Pro | 58.6% | 64.3% | Claude Opus 4.7 |
| Terminal-Bench 2.0 | 82.7% | 69.4% | GPT-5.5 |
| Humanity's Last Exam | 40.6% | 31.2% | GPT-5.5 |
| Humanity's Last Exam with tools | 52.2% | 54.7% | Claude Opus 4.7 |
| BrowseComp | 84.4% | 79.3% | GPT-5.5 |
| GPQA Diamond | 93.6% | 94.2% | Claude Opus 4.7 |
| ARC-AGI-1 Verified | 94.5% | 92.0% | GPT-5.5 |
On those numbers, Claude looks better on SWE-Bench Pro, GPQA Diamond and Humanity's Last Exam with tools, while GPT-5.5 looks better on Terminal-Bench 2.0, BrowseComp, Humanity's Last Exam without tools and ARC-AGI-1 Verified.
LLM Stats' common 10-benchmark view comes out 6-4 for Claude Opus 4.7. Its interpretation is more useful than the raw tally: Opus 4.7 is stronger on reasoning-heavy and review-grade tests, while GPT-5.5 is stronger on long-running tool-use tests.
There is an important caveat. LLM Stats notes that the compared scores are self-reported at each provider's high-reasoning tier, making them similar in shape but not identical in methodology. Even a benchmark such as Humanity's Last Exam can look different depending on the source and exact setup.
Kimi K2.6 and DeepSeek V4 Pro are easiest to evaluate as open-weight operating choices rather than direct one-for-one replacements for closed frontier models. Using the Artificial Analysis open-model table, the trade-off is straightforward.
| Metric | Kimi K2.6 | DeepSeek V4 Pro |
|---|---|---|
| Artificial Analysis Intelligence | 54 | 52 |
| Context window | 256k | 1.00M |
| Price column | $1.7 | $2.2 |
| Output speed | 112 tokens/s | 36 tokens/s |
Taken at face value, Kimi K2.6 has the edge on Intelligence score and output speed, while DeepSeek V4 Pro has the much larger context window. The Decoder also reports Moonshot AI's published Kimi K2.6 figures of 54.0 on HLE with Tools, 58.6 on SWE-Bench Pro and 83.2 on BrowseComp.
But Kimi's public testing should not be read as a perfectly controlled comparison with GPT-5.5 and Claude Opus 4.7. The Hugging Face model card says Kimi K2.6 was evaluated with thinking mode enabled, temperature 1.0, top-p 1.0 and a 262,144-token context length, and its listed comparisons are mainly against Claude Opus 4.6, GPT-5.4 and Gemini 3.1 Pro.
DeepSeek V4 Pro, meanwhile, is less about winning every hardest benchmark and more about long context and cost. DataCamp summarizes DeepSeek V4 as not surpassing GPT-5.5 or Claude Opus 4.7 in pure capability, but as targeting near-frontier performance at lower cost.
If a model is going into production, price is not one number. Keep at least three cost measures separate.
API token pricing. Mashable reports DeepSeek V4 at $1.74 per 1M input tokens and $3.48 per 1M output tokens, GPT-5.5 at $5 and $30, and Claude Opus 4.7 at $5 and $25.
Artificial Analysis Price column. The open-model table lists Kimi K2.6 at $1.7 and DeepSeek V4 Pro at $2.2, but that column should not be treated as identical to Mashable's API token prices.
Benchmark execution cost. Artificial Analysis says DeepSeek V4 Pro costs $1,071 to run the Intelligence Index, compared with $948 for Kimi K2.6 and $4,811 for Claude Opus 4.7.
So a claim such as DeepSeek is cheaper, Kimi is cheaper or Claude is expensive is only meaningful after you specify the meter: API input/output tokens, the Artificial Analysis price column, benchmark run cost, or your own workload's real output-token volume.
Benchmarks such as SWE-Bench Pro, GPQA Diamond and BrowseComp measure task performance. They do not automatically settle questions about honesty, sycophancy, hallucination risk or auditability.
For Claude Opus 4.7, Mashable reports Anthropic's claim of a 92% honesty rate and less sycophancy. Anthropic also says Claude Opus 4.7 tied for the top overall score across six modules on its internal research-agent benchmark at 0.715, and improved on General Finance from 0.767 for Opus 4.6 to 0.813.
Those are useful signals, but they belong on a different axis from coding, science Q&A, web browsing and terminal automation scores. In production, evaluate capability, cost, speed, failure modes and reviewability separately.
For many teams, the practical answer is not to pick one model forever. It is to route tasks. MindStudio reports that GPT-5.5 used 72% fewer output tokens than Claude Opus 4.7 on the same coding tasks, while also arguing that Opus 4.7's thoroughness can justify the cost for complex, reasoning-heavy work across large codebases.
A sensible starting setup would be: GPT-5.5 for standard generation, coding fixes, terminal and browser-heavy work; Claude Opus 4.7 for deep review and high-stakes reasoning; Kimi K2.6 for lower-cost open-weight experiments; and DeepSeek V4 Pro for long-context or high-volume tasks where DeepSeek V4's API pricing matters.
Based on the public evidence, the safest conclusion is not that one model wins everything. GPT-5.5 has the strongest broad and economic-task signals, Claude Opus 4.7 is especially compelling for reasoning and review, Kimi K2.6 is attractive for open-weight speed and price/performance, and DeepSeek V4 Pro stands out for long context and lower DeepSeek V4 API token pricing.
Even Artificial Analysis can show different pictures depending on page and setting: its GPT-5.5 model page lists GPT-5.5 high at Intelligence 59, while another Artificial Analysis models page says Claude Opus 4.7, Adaptive Reasoning, Max Effort leads the Intelligence Index at 57.
Use benchmarks as a shortlist generator, not as a procurement decision by themselves. The final call should come from small parallel tests on your own prompts, budget, latency targets and tolerance for errors.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Treat the scores as use case signals, not a final league table. Public results mix reasoning settings, reporting dates, provider reported numbers and third party analysis.
Treat the scores as use case signals, not a final league table. Public results mix reasoning settings, reporting dates, provider reported numbers and third party analysis. GPT 5.5 has strong broad and economic task signals: Artificial Analysis lists GPT 5.5 high at 59 on its Intelligence Index and GPT 5.5 xhigh at Elo 1785 on GDPval AA.
Claude Opus 4.7 leads GPT 5.5 on 6 of 10 common benchmarks in LLM Stats, while Kimi K2.6 stands out for open weight speed and DeepSeek V4 Pro for 1M context and lower DeepSeek V4 API token pricing.