Google喺2026年7月8日更新Android Bench,Anthropic嘅Claude Fable 5以84.5分登上榜首,同時全面改用開源Harbor框架,令所有舊分數無法直接比較。 呢次更新首次引入開發者貢獻機制,容許開發者提交真實Android開發任務,以及分享自己用Harbor工具做嘅模型評測,等大家可以直接比拼。
研究答案

Create a landscape editorial hero image for this Studio Global article: Search & fact-check with cited sources for What does Google's July 8, 2026 update to Android Bench reveal about the new top-ranked AI model. Article summary: On **July 8, 2026**, Google issued a major update to **Android Bench** — its official leaderboard for evaluating LLMs on real-world Android development tasks. The update includes a new #1 ranked model, a switch to the op. Topic tags: general, documentation, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
2026年7月8日,Google正式發布Android Bench嘅重大更新D9。Android Bench係Google官方用來評估大型語言模型(LLM)喺Android開發任務上表現嘅排行榜。呢次更新帶嚟咗全新嘅冠軍模型、全面改用開源Harbor框架、首次開放社群貢獻,仲有一張反映成本同準確度之間殘酷取捨嘅全新排行榜DA。
Anthropic嘅Claude Fable 5以84.5分成為Android Bench新霸主,領先第二名超過4分DA。呢個分數反映咗模型喺100個真實Android編程任務嘅表現,呢啲任務係從大約39,000個開源pull request入面揀出嚟嘅DG。
Fable 5喺其他業界編程指標都係頂尖水平:
多個獨立排行榜喺2026年7月都將Fable 5列為整體最強AI模型BB。
所有之前上榜嘅模型都係用新嘅Harbor方法重新評分,所以得出咗一張煥然一新嘅排行榜DA。以下係Android Bench前十名:
| 排名 | 模型 | 分數 | 平均延遲(秒) | 平均成本($/1K任務) |
|---|---|---|---|---|
| 1 | Claude Fable 5(Anthropic) | 84.5 | 8.0 | $133.20 |
| 2 | GPT 5.5(OpenAI) | 80.2 | 15.7 | $138.30 |
| 3 | Claude Sonnet 5(Anthropic) | 76.2 | 12.3 | $99.90 |
| 4 | GPT 5.4(OpenAI) | 74.1 | 8.4 | $83.40 |
| 5 | Gemini 3.1 Pro Preview(Google) | 73.7 | 10.6 | $87.40 |
| 6 | Claude Opus 4.8(Anthropic) | 72.4 | 6.7 | $88.00 |
| 7 | GLM 5.2 | 72.2 | 38.9 | $117.00 |
| 8 | Gemini 3.5 Flash(Google) | 71.1 | 28.3 | $165.60 |
| 9 | Kimi K2.7 Code | 70.4 | 31.8 | $48.10 |
| 10 | (省略) | — | — | — |
來源:9to5Google報導Android Bench更新後嘅排名A
值得留意嘅幾點:
Google將Android Bench標準化咗喺Harbor框架上面。Harbor係由Laude Institute(即Terminal-Bench嘅開發團隊)開發嘅開源評估生態系統DATG。之前Android Bench用嘅係自家開發嘅mini-swe-agent v1測試工具。Harbor提供咗一個標準化、基於容器嘅評估流程,支援雲端部署、社群任務提交同強化學習訓練ATQ。
呢個轉變意味住所有之前嘅模型分數都冇得直接比較——上面列出嘅每個分數都係用Harbor方法重新評估嘅結果9A。Google話咁做係為咗令評估標準跟得上LLM嘅快速發展D。
Google第一次開放Android Bench畀社群貢獻9A。開發者而家可以:
Google話呢個係回應開發者要求「提供一個渠道反饋我哋嘅數據集」,係同Android開發社群更深層合作嘅一步9。
更新後嘅排行榜揭示咗一個充滿取捨嘅市場A:
2026年7月嘅更新再次確認咗Anthropic、OpenAI同Google之間嘅三強競爭A:
今次係Anthropic嘅型號首次喺Google自己嘅Android編程排行榜上登頂,喺手機AI助手市場上有標誌性嘅意義。
Studio Global AI
此頁麵包含一個有來源支援的答案,您可以在 Studio Global 內繼續。
Google喺2026年7月8日更新Android Bench,Anthropic嘅Claude Fable 5以84.5分登上榜首,同時全面改用開源Harbor框架,令所有舊分數無法直接比較。
Google喺2026年7月8日更新Android Bench,Anthropic嘅Claude Fable 5以84.5分登上榜首,同時全面改用開源Harbor框架,令所有舊分數無法直接比較。 呢次更新首次引入開發者貢獻機制,容許開發者提交真實Android開發任務,以及分享自己用Harbor工具做嘅模型評測,等大家可以直接比拼。
成本分析顯示殘酷現實:冠軍Claude Fable 5每1K任務成本133.2美元,而最平嘅Kimi K2.7 Code只係48.1美元,但準確度低成14.1分。Gemini 3.5 Flash更係又貴又慢又唔準。