| 終端機操作、瀏覽、長時間工具使用 | GPT-5.5 | LLM Stats 指出 GPT-5.5 在 Terminal-Bench 2.0、BrowseComp、OSWorld-Verified、CyberGym 等長時間工具使用測試較強。 |
| 開放權重路線,重視速度與價格性能 | Kimi K2.6 | Artificial Analysis 開放模型表列出 Kimi K2.6:Intelligence 54、256k context、Price 欄位 $1.7、112 tokens/s。 |
| 長上下文、大量處理、低 API 單價 | DeepSeek V4 Pro/DeepSeek V4 系列 | Artificial Analysis 顯示 DeepSeek V4 Pro 有 1M context;Mashable 報告 DeepSeek V4 的 API 單價低於 GPT-5.5 與 Claude Opus 4.7。 |
GPT-5.5 與 Claude Opus 4.7 都是前沿閉源模型,但勝負會隨基準測試而變。以 Mashable 報告的數字看,Claude Opus 4.7 在 SWE-Bench Pro 與 GPQA Diamond 領先;GPT-5.5 則在 Terminal-Bench 2.0、Humanity's Last Exam、BrowseComp、ARC-AGI-1 Verified 領先。
| 基準測試 | GPT-5.5 | Claude Opus 4.7 | Mashable 表中領先者 |
|---|---|---|---|
| SWE-Bench Pro | 58.6% | 64.3% | Claude Opus 4.7 |
| Terminal-Bench 2.0 | 82.7% | 69.4% | GPT-5.5 |
| Humanity's Last Exam | 40.6% | 31.2% | GPT-5.5 |
| Humanity's Last Exam with tools | 52.2% | 54.7% | Claude Opus 4.7 |
| BrowseComp | 84.4% | 79.3% | GPT-5.5 |
| GPQA Diamond | 93.6% | 94.2% | Claude Opus 4.7 |
| ARC-AGI-1 Verified | 94.5% | 92.0% | GPT-5.5 |
LLM Stats 的歸納則是:在雙方都有回報的 10 項基準中,Claude Opus 4.7 領先 6 項,GPT-5.5 領先 4 項;Opus 4.7 偏強於推理、審查與專業任務,GPT-5.5 偏強於長時間工具使用。
但這裡要特別小心。LLM Stats 也提醒,這些分數多來自各供應商高推理層級的自報結果,形式上可以對照,方法論卻未必完全相同。 甚至像 Humanity's Last Exam 這類項目,不同來源呈現出的領先方向也可能不同。
Kimi K2.6 與 DeepSeek V4 Pro 不宜直接拿來和 GPT-5.5、Claude Opus 4.7 做單一總分競賽;更實際的角度,是把它們視為開放權重部署或實驗的候選。
| 指標 | Kimi K2.6 | DeepSeek V4 Pro |
|---|---|---|
| Artificial Analysis Intelligence | 54 | 52 |
| Context window | 256k | 1.00M |
| Price 欄位 | $1.7 | $2.2 |
| Output speed | 112 tokens/s | 36 tokens/s |
只看這張表,Kimi K2.6 在 Intelligence 與輸出速度上較有利;DeepSeek V4 Pro 的明顯優勢則是 1M context。 The Decoder 也轉述 Moonshot AI 發表值,稱 Kimi K2.6 在 HLE with Tools 為 54.0、SWE-Bench Pro 為 58.6、BrowseComp 為 83.2。
不過,Kimi K2.6 的公開實驗並不是與 GPT-5.5、Claude Opus 4.7 完全同條件對打。Hugging Face 模型卡說明,Kimi K2.6 以 thinking mode、temperature 1.0、top-p 1.0、262,144 token 上下文長度等設定評估,主要比較對象也包括 Claude Opus 4.6、GPT-5.4、Gemini 3.1 Pro。
DeepSeek V4 Pro 則更像是用長上下文與成本換取接近前沿模型能力的方案,而不是純性能冠軍。DataCamp 整理指出,DeepSeek V4 在純能力上沒有超過 GPT-5.5 與 Claude Opus 4.7,但定位是以較低成本提供 near-frontier 性能。
看價格時,至少要分清三種數字。
第一是 API token 單價。Mashable 報告 DeepSeek V4 為每 100 萬輸入 token $1.74、每 100 萬輸出 token $3.48;GPT-5.5 為 $5/$30;Claude Opus 4.7 為 $5/$25。
第二是 Artificial Analysis 模型表中的 Price 欄位。該表列出 Kimi K2.6 為 $1.7、DeepSeek V4 Pro 為 $2.2,但這不應直接等同於 Mashable 報告的 API token 單價。
第三是跑完整個基準測試的成本。Artificial Analysis 文章指出,執行 Intelligence Index 的成本為 DeepSeek V4 Pro $1,071、Kimi K2.6 $948、Claude Opus 4.7 $4,811。
Claude Opus 4.7 的安全與可靠性訊號值得另外看。Mashable 轉述 Anthropic 說法,稱 Claude Opus 4.7 有 92% honesty rate,且 sycophancy 較少。 Anthropic 自家發布也表示,Claude Opus 4.7 在內部 research-agent benchmark 的 6 個模組總分並列第一,達 0.715;在 General Finance 模組中,分數由 Opus 4.6 的 0.767 提升到 0.813。
如果是生產環境,硬把所有任務交給同一個模型,通常不是最穩的做法。MindStudio 的程式任務比較指出,GPT-5.5 在相同 coding task 中比 Claude Opus 4.7 少用 72% 輸出 token;但對複雜、推理負荷高的大型程式碼庫,Opus 4.7 的細緻程度可能足以支撐較高成本。
較務實的配置是:標準生成、修改、終端機與工具型任務先試 GPT-5.5;深度審查、專業判斷與高風險推理交給 Claude Opus 4.7;開放權重與低成本實驗測 Kimi K2.6;長上下文、大量文件或低 API 單價需求則測 DeepSeek V4 Pro。
目前公開資料更支持用途分流,而不是宣布單一總冠軍。GPT-5.5 的亮點在綜合能力與經濟價值任務;Claude Opus 4.7 適合推理、審稿與專業判斷;Kimi K2.6 是開放權重路線中速度與價格性能突出的選項;DeepSeek V4 Pro 則以長上下文與 DeepSeek V4 系列較低 API 單價吸引人。