A dash means the supplied source excerpt did not include a comparable result for that model.
The biggest trap is treating unlike rows as if they were one unified leaderboard. Artificial Analysis compares GPT-5.5 medium, Kimi K2.6 and Claude Opus 4.7 non-reasoning high; AkitaOnRails uses GPT-5.5 xHigh/Codex and separate DeepSeek V4 Flash and Pro rows; VentureBeat also separates GPT-5.5 from GPT-5.5 Pro.
Even the direct GPT-5.5 versus Claude Opus 4.7 picture is mixed. LLM Stats says that across 10 benchmarks reported by both providers, Opus 4.7 leads on six, while GPT-5.5 leads on four. Claude’s wins cluster around reasoning-heavy and review-grade tests, while GPT-5.5’s wins cluster around long-running tool use and shell-driven tasks.
GPT-5.5’s best cited signals are ARC and Terminal-Bench. In DocsBot’s ARC comparison, GPT-5.5 scores 85% on ARC-AGI-2 versus 75.8% for Claude Opus 4.7, and 95% on ARC-AGI-1 versus 93.5% for Claude. In VentureBeat’s Terminal-Bench 2.0 row, GPT-5.5 reaches 82.7%, well above Claude Opus 4.7 at 69.4% and DeepSeek at 67.9%.
Artificial Analysis also places GPT-5.5 medium above the two directly visible competitors in that excerpt: 57 for GPT-5.5 medium, 54 for Kimi K2.6 and 52 for Claude Opus 4.7 non-reasoning high. That should not be read as a universal win across every Claude or Kimi mode, but it is a useful snapshot.
Claude Opus 4.7 is most compelling in the cited hard-reasoning and software-engineering evaluations. On Humanity’s Last Exam without tools, VentureBeat lists Claude at 46.9%, GPT-5.5 at 41.4% and DeepSeek at 37.7%. With tools enabled, Claude is at 54.7%, GPT-5.5 at 52.2% and DeepSeek at 48.2%, though the separate GPT-5.5 Pro row is higher at 57.2%.
On SWE-Bench Pro, DataCamp lists Claude Opus 4.7 at 64.3%, GPT-5.5 at 58.6% and DeepSeek V4 Pro at 55.4%. That lines up with LLM Stats’ broader summary: Claude leads GPT-5.5 on GPQA, HLE without tools, HLE with tools, SWE-Bench Pro, MCP Atlas and FinanceAgent v1.1, while GPT-5.5 leads on Terminal-Bench 2.0, BrowseComp, OSWorld-Verified and CyberGym.
Kimi K2.6 is harder to rank across all four models because it is not present in every shared benchmark table. In the Artificial Analysis excerpt, Kimi K2.6 scores 54, below GPT-5.5 medium at 57 but above Claude Opus 4.7 non-reasoning high at 52.
In AkitaOnRails’ coding benchmark, Kimi K2.6 scores 87. That is below Claude Opus 4.7 at 97 and GPT-5.5 xHigh/Codex at 96, but above DeepSeek V4 Flash at 78 and DeepSeek V4 Pro at 69. Verdent separately lists SWE-Bench Verified at 80.2% for Kimi K2.6 and 87.6% for Claude Opus 4.7.
Kimi’s practical differentiator is the open-weight route. Verdent says K2.6 weights are available on Hugging Face and can run on vLLM, SGLang or KTransformers, with a minimum viable setup of 4× H100 GPUs for the INT4 variant at reduced context. A Hugging Face README also lists Kimi K2.6 agentic metrics such as HLE-Full with tools at 54.0, BrowseComp at 83.2, DeepSearchQA f1-score at 92.5, Toolathlon at 50.0 and MCPMark at 55.9, but that table mostly compares Kimi with GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro rather than the full four-model set here.
In the cited benchmark rows, DeepSeek V4 is usually not the maximum-score option. VentureBeat places DeepSeek below GPT-5.5 and Claude Opus 4.7 on Humanity’s Last Exam without tools, Humanity’s Last Exam with tools and Terminal-Bench 2.0. DataCamp lists DeepSeek V4 Pro at 55.4% on SWE-Bench Pro, behind GPT-5.5 at 58.6% and Claude Opus 4.7 at 64.3%.
AkitaOnRails lists DeepSeek V4 Flash at 78 and DeepSeek V4 Pro at 69, below Kimi K2.6, GPT-5.5 xHigh/Codex and Claude Opus 4.7 in the same table.
The pricing, however, is materially different. Mashable lists DeepSeek V4 at $1.74 per 1 million input tokens and $3.48 per 1 million output tokens. The same comparison lists GPT-5.5 at $5 per 1 million input tokens and $30 per 1 million output tokens, and Claude Opus 4.7 at $5 and $25 respectively. That does not make DeepSeek V4 the benchmark winner, but it can make it attractive for high-volume drafts, low-risk experiments and internal evaluations where cost per attempt matters more than the top score.
On benchmarks alone, the top tier in these sources is GPT-5.5 and Claude Opus 4.7, but they win in different places. GPT-5.5 is stronger in ARC and Terminal-Bench, while Claude Opus 4.7 is stronger in HLE and SWE-Bench Pro. Kimi K2.6 remains a credible coding and agentic model, especially where an open-weight deployment path matters, but there are fewer direct shared comparisons.
DeepSeek V4 is generally lower on the cited raw scores, yet its API price makes it a serious candidate for price-performance pilots.