Fable 5 also leads broader coding benchmarks industry-wide. Independent reporting shows it scoring 95.0% on SWE-bench Verified and 80.0% on SWE-bench Pro, well ahead of prior-generation models WLM. On the agentic coding benchmark Terminal-Bench 2.0, Fable 5 recorded an 84.3% accuracy WB. Multiple independent leaderboards now rank Fable 5 as the top overall AI model as of July 2026 BB.
All previously ranked models were reevaluated under the new Harbor-based methodology, producing a substantially refreshed leaderboard DA. The top ten according to Android Bench:
| Rank | Model | Score | Avg Latency (s) | Avg Cost ($/1K tasks) |
|---|---|---|---|---|
| 1 | Claude Fable 5 (Anthropic) | 84.5 | 8.0 | $133.20 |
| 2 | GPT 5.5 (OpenAI) | 80.2 | 15.7 | $138.30 |
| 3 | Claude Sonnet 5 (Anthropic) | 76.2 | 12.3 | $99.90 |
| 4 | GPT 5.4 (OpenAI) | 74.1 | 8.4 | $83.40 |
| 5 | Gemini 3.1 Pro Preview (Google) | 73.7 | 10.6 | $87.40 |
| 6 | Claude Opus 4.8 (Anthropic) | 72.4 | 6.7 | $88.00 |
| 7 | GLM 5.2 | 72.2 | 38.9 | $117.00 |
| 8 | Gemini 3.5 Flash (Google) | 71.1 | 28.3 | $165.60 |
| 9 | Kimi K2.7 Code | 70.4 | 31.8 | $48.10 |
Source: 9to5Google's reporting from Android Bench's refreshed rankings A.
Key observations from the updated standings:
Google standardized Android Bench on the Harbor framework, an open-source evaluation ecosystem developed by the Laude Institute, the team behind Terminal-Bench DATG. Previously, Android Bench used a custom mini-swe-agent v1 harness. Harbor provides a standardized, container-based evaluation pipeline that supports cloud deployment, community task submission, and reinforcement learning rollouts ATQ.
The migration means all prior model scores are non-comparable — every score listed above is a fresh evaluation under the Harbor methodology 9A. Google stated the move was necessary to keep evaluation standards state-of-the-art as LLMs rapidly improve D.
For the first time, Google has opened Android Bench to community contributions 9A. Developers can now:
Google described this as a response to developer demand for "a way to provide feedback on our dataset" and a move toward deeper collaboration with the Android development community 9.
The refreshed leaderboard reveals a market with significant spread between cost leaders and accuracy leaders A:
The July 2026 update reinforces a three-way race among Anthropic, OpenAI, and Google A:
This is the first Android Bench update where an Anthropic model leads Google's own benchmark for Android coding, a symbolic shift in the mobile AI assistant market.