公开数据不足以支持一个绝对“总冠军”:GPT 5.5 在可见 Intelligence Index、BrowseComp 和 Terminal Bench 2.0 上突出;Claude Opus 4.7 在 GPQA Diamond 与 Humanity’s Last Exam no tools 上领先;Kimi K2.6 缺少完整四方同场数据。[2][7][4] DeepSeek V4 的最大优势是成本:公开摘要列出其每 100 万 token 输入 / 输出价格为 1.74 / 3.48 美元,低于 GPT 5.5 的 5 / 30 美元与 Claude Opus 4.7 的 5 / 25 美元。[1][17] 实务选型更...

Create a landscape editorial hero image for this Studio Global article: GPT-5.5、Claude Opus 4.7、DeepSeek V4、Kimi K2.6 怎麼選?Benchmark 與價格比較. Article summary: 公開數據不支持一個絕對總冠軍:GPT 5.5 在可見 Intelligence Index 60/59、BrowseComp 84.4% 與 Terminal Bench 2.0 82.7% 最突出;Claude Opus 4.7 在 GPQA Diamond 94.2% 與 HLE no tools 46.9% 領先,Kimi K2.6 則缺少完整四方同場數據。[2][7]. Topic tags: ai, llm benchmarks, openai, anthropic, deepseek. Reference image context from search candidates: Reference image 1: visual subject "[Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison](https://www.youtube.com/watch?v=M90iB4hpenI). . [](https://www.youtube.com" source context "Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison - YouTube" Reference image 2: visual subject "[Kimi K2 vs Claude Opus 4.7 vs GPT 5.5 Comparison](https://www.youtube.com/watch?v=M90iB4hpenI). 。LLM Stats 也提醒,GPT-5.5 与 Claude Opus 4.7 的部分分数是供应商在高推理 tier 下自报,趋势可以参考,但方法论并不完全一致。
所以,更可靠的读法不是问“哪个模型最强”,而是先问:你要解决什么任务?工具型代理看 GPT-5.5,推理与审查看 Claude Opus 4.7,成本敏感 API 看 DeepSeek V4,开源 coding-agent 探索则把 Kimi K2.6 放进实测清单。
DeepSeek 的公开资料口径并不完全统一:价格来源多写 DeepSeek V4 或 DeepSeek V4 Pro,部分 benchmark 则写 DeepSeek-V4-Pro-Max。 下表保留来源中的名称,避免把不同设置误读为完全相同的模型配置。
Artificial Analysis 的可见摘要列出 Intelligence Index 前五名:GPT-5.5 xhigh 为 60、GPT-5.5 high 为 59、Claude Opus 4.7 Adaptive Reasoning, Max Effort 为 57,后面还有 Gemini 3.1 Pro Preview 与 GPT-5.4 xhigh,同为 57。
这能支持一个有限结论:在该摘要可见的 Intelligence Index 领先模型中,GPT-5.5 排在 Claude Opus 4.7 前面。 但它不能直接推出四款模型的完整总排名,因为同一可见摘要没有给出 DeepSeek V4 与 Kimi K2.6 的同口径 Intelligence Index 分数。
BrowseComp 偏向评估 agentic AI web browsing,尤其是高度容器化的信息查找任务。VentureBeat 摘要列出的结果是:GPT-5.5 84.4%、DeepSeek-V4-Pro-Max 83.4%、Claude Opus 4.7 79.3%。
这意味着,在 web browsing 代理任务上,DeepSeek-V4-Pro-Max 与 GPT-5.5 的差距很小;但 Claude Opus 4.7 在同表中落后一些。
Terminal-Bench 2.0 的差距更明显。VentureBeat 摘要列出 GPT-5.5 82.7%、Claude Opus 4.7 69.4%、DeepSeek 67.9%;Yahoo / Investing.com 也描述 Terminal-Bench 2.0 测试 command-line workflows,并列出 GPT-5.5 82.7%。
Kimi K2.6 的 Terminal-Bench 2.0 可见数字为 66.70%,但来源比较的是 Kimi K2.6、Claude Opus 4.6 与 GPT-5.4,不是 GPT-5.5、Claude Opus 4.7、DeepSeek V4 的同场表。
VentureBeat 摘要列出 GPQA Diamond:Claude Opus 4.7 94.2%、GPT-5.5 93.6%、DeepSeek-V4-Pro-Max 90.1%。同一摘要列出 Humanity’s Last Exam no-tools:Claude Opus 4.7 46.9%、GPT-5.5 41.4%、GPT-5.5 Pro 43.1%、DeepSeek-V4-Pro-Max 37.7%。
LLM Stats 对 GPT-5.5 与 Claude Opus 4.7 的结论也指向同一方向:在双方都报告的 10 个 benchmark 中,Claude Opus 4.7 领先 6 个,GPT-5.5 领先 4 个;Claude 的优势集中在 reasoning-heavy 与 review-grade tests,而 GPT-5.5 的优势集中在 long-running tool-use tests。
如果你的任务是复杂推理、合规审查、财务分析复核、代码审查或低容错决策,Claude Opus 4.7 值得优先进入测试队列。
DataCamp 的 DeepSeek V4 对比表列出 SWE-Bench Pro:DeepSeek V4 Pro 55.4%、GPT-5.5 58.6%、Claude Opus 4.7 64.3%。 Yahoo / Investing.com 也称 GPT-5.5 在 SWE-Bench Pro 为 58.6%,并说明该测试评估 GitHub issue resolution。
Kimi K2.6 的 coding 数字值得单独看。Verdent 摘要列出 Kimi K2.6 在 SWE-Bench Pro 为 58.60%、SWE-Bench Verified 为 80.20%、LiveCodeBench v6 为 89.60%;但同一摘要注明,Kimi K2.6 数字来源为 Moonshot AI official model card,且 SWE-Bench Pro 使用 Moonshot in-house harness。
因此,Kimi K2.6 可以列入 coding-agent 候选,但不适合直接拿这些数字硬排进四方总榜。
更实际的做法是:如果任务是大型 repo 修复、code review 或长时间 coding agent,不要只看单一 SWE 分数。Claude Opus 4.7 在可见 SWE-Bench Pro 对比中最高;GPT-5.5 在 Terminal-Bench 2.0 这类长流程工具任务上领先;Kimi K2.6 则需要用自己的 repo、CI 流程和工具链补测。
Mashable 摘要列出三款模型的 API 价格:DeepSeek V4 为每 100 万输入 token 1.74 美元、每 100 万输出 token 3.48 美元,并标示 1M context window;GPT-5.5 为每 100 万输入 5 美元、输出 30 美元,并标示 1M context window;Claude Opus 4.7 为每 100 万输入 5 美元、输出 25 美元,并标示 1M context window。
DataCamp 的 DeepSeek V4 对比摘要也使用相同价格口径,并列出 DeepSeek V4 Pro、GPT-5.5、Claude Opus 4.7 的 context window 约为 1M tokens。
在这些可见价格中,DeepSeek V4 明显低于 GPT-5.5 与 Claude Opus 4.7;再加上 DeepSeek-V4-Pro-Max 在 BrowseComp 为 83.4%、接近 GPT-5.5 的 84.4%,它很适合作为成本敏感 API 路由的第一批测试对象。
Kimi K2.6 的同口径 API 价格没有出现在提供来源中;DocsBot 摘要则称 Kimi K2.6 具 256K context,并将其描述为面向 long-horizon coding、coding-driven design、autonomous execution 与 swarm-based orchestration 的 open-source agentic model。
对多数产品团队来说,最务实的答案不是“只买哪一个模型”,而是先建立分层路由与回归测试:
如果只用可见公开资料初筛,GPT-5.5 是 agentic tool-use 与可见综合排名的最强候选;Claude Opus 4.7 是推理与 review-grade 任务的最强候选之一;DeepSeek V4 是价格最有吸引力的高性价比候选;Kimi K2.6 则应放进开源 / coding-agent 实验池,但目前证据不足以公平排入完整四方总榜。
上线或采购前,建议用同一批真实任务做回归测试:同一 prompt、同一工具权限、同一上下文长度、同一成功判准。公开 benchmark 的价值,是帮你决定先测谁;最终选型,仍应由你的产品场景、错误成本与 token 成本共同决定。
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
公开数据不足以支持一个绝对“总冠军”:GPT 5.5 在可见 Intelligence Index、BrowseComp 和 Terminal Bench 2.0 上突出;Claude Opus 4.7 在 GPQA Diamond 与 Humanity’s Last Exam no tools 上领先;Kimi K2.6 缺少完整四方同场数据。[2][7][4]
公开数据不足以支持一个绝对“总冠军”:GPT 5.5 在可见 Intelligence Index、BrowseComp 和 Terminal Bench 2.0 上突出;Claude Opus 4.7 在 GPQA Diamond 与 Humanity’s Last Exam no tools 上领先;Kimi K2.6 缺少完整四方同场数据。[2][7][4] DeepSeek V4 的最大优势是成本:公开摘要列出其每 100 万 token 输入 / 输出价格为 1.74 / 3.48 美元,低于 GPT 5.5 的 5 / 30 美元与 Claude Opus 4.7 的 5 / 25 美元。[1][17]
实务选型更适合按任务路由:GPT 5.5 先测工具代理与浏览,Claude Opus 4.7 先测推理与审查,DeepSeek V4 先测高流量 API,Kimi K2.6 放进开源 coding agent 实验池。[3][5][7]