Empty cells do not mean DeepSeek V4 or Kimi K2.6 performed poorly. They mean this evidence set does not include matched scores on the same tests, with the same settings and the same level of detail
.
On the two ARC-AGI scores published by OpenAI, GPT-5.5 beats Claude Opus 4.7: 95.0% versus 93.5% on ARC-AGI-1 Verified, and 85.0% versus 75.8% on ARC-AGI-2 Verified .
That is the cleanest reasoning comparison in the evidence set. It is not, however, a universal verdict. OpenAI notes that GPT evaluations were run with reasoning effort set to xhigh in a research environment, which may produce outputs slightly different from production ChatGPT in some cases . For buyers and builders, that caveat matters: a benchmark run is not always the same as the API behavior you will see in a live product.
The strongest cited result for Claude Opus 4.7 comes from MCP-Atlas. A secondary analysis reports Claude Opus 4.7 at 79.1% versus GPT-5.5 at 75.3%, linking Claude's lead to better tool-call reliability in complex, chained scenarios via the Model Context Protocol .
That may matter as much as abstract reasoning for teams building multi-tool agents. If your product depends on reliable tool calls, external systems and chained workflows, the best cited benchmark signal here favors Claude Opus 4.7 over GPT-5.5 on MCP-Atlas specifically .
GPT-5.5 is reported at 82.7% on Terminal-Bench 2.0, a benchmark tied to terminal tasks and agentic coding . Among the sources cited for this comparison, that is the clearest numerical coding signal.
The limitation is just as important as the score. The cited sources do not provide a complete Terminal-Bench 2.0 grid for Claude Opus 4.7, DeepSeek V4 and Kimi K2.6. A careful conclusion is therefore narrower: GPT-5.5 has the strongest documented signal here, but the evidence does not prove it beats all three alternatives under every agentic-coding setup .
DeepSeek V4 and Kimi K2.6 should not be dismissed. They matter because open-weights models can be attractive when teams want more control over deployment, tuning or infrastructure choices. But the cited sources do not provide matched ARC-AGI, MCP-Atlas or Terminal-Bench 2.0 scores for a rigorous four-way comparison
.
For DeepSeek, Artificial Analysis says the release of DeepSeek V4 brings DeepSeek back among the leading open-weights models . The most precise figure supplied here is for DeepSeek V4 Pro (Max): 52 on the Artificial Analysis Intelligence Index, up from 42 for DeepSeek V3.2
.
For Kimi K2.6, Artificial Analysis highlights an analysis titled Kimi K2.6: The new leading open weights model . That is a strong positioning signal, but it is not the same as a shared benchmark table showing Kimi K2.6 against GPT-5.5, Claude Opus 4.7 and DeepSeek V4 on the same tests
.
Safety evidence needs its own lane. GPT-5.5's system card describes CoT-Control as a suite of more than 13,000 tasks built from established benchmarks including GPQA, MMLU-Pro, HLE, BFCL and SWE-Bench Verified . That helps explain how OpenAI evaluates controllability of reasoning behavior, but it does not rank GPT-5.5 against Claude Opus 4.7, DeepSeek V4 and Kimi K2.6
.
A separate source reports a 93% cyber range pass rate for GPT-5.5 while also reporting that a universal jailbreak was found in six hours of red-teaming . Read together, those claims underline the point: strong cyber-task performance is not the same as global model safety
.
An external critique also argues that GPT-5.5 safety assessment still depends heavily on OpenAI's own claims, limiting what can be concluded from supplier-published information alone . That does not invalidate the benchmark results, but it does mean safety-sensitive deployments need more than headline scores.
xhigh reasoning effort in a research environment Do not conclude that GPT-5.5 is the universal best model just because it leads Claude Opus 4.7 on the ARC-AGI scores available here . Do not conclude that Claude Opus 4.7 is globally superior just because it wins on MCP-Atlas
. Those benchmarks measure different things.
Do not force DeepSeek V4 and Kimi K2.6 into a four-way ranking without shared benchmark data. The Artificial Analysis signals show that both models are important in the open-weights ecosystem, but they do not establish a clean global leaderboard against GPT-5.5 and Claude Opus 4.7 on the same metrics
.
And do not treat a capability score as a safety guarantee. The available GPT-5.5 reporting shows exactly why: strong cyber performance can coexist with jailbreak concerns and questions about the independence of safety evaluation
.
The most honest ranking is by use case, not by hype. GPT-5.5 leads the cited ARC-AGI comparisons against Claude Opus 4.7 and has the clearest cited signal for agentic coding. Claude Opus 4.7 leads on MCP-Atlas. DeepSeek V4 and Kimi K2.6 remain important open-weights contenders, but the available sources do not rank them cleanly against the two proprietary models on the same benchmark set
.
For a product decision, the right next step is not to crown a universal winner. It is to run your own evaluation on the tasks that matter: reasoning, tool calls, code, cost, latency, deployment constraints and acceptable risk.