Google overhauls its coding-agent benchmark
Archive item — written before sources were shown.
Android Bench switched to the Harbor evaluation framework and added eight models; Claude Fable 5 currently leads at 84.5 versus GPT-5.5's 80.2.
Google overhauled its Android Bench coding-agent benchmark on July 8, switching its evaluation harness to the standardized “Harbor” framework and adding eight new models, including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, and two Qwen variants. Claude Fable 5 currently leads the refreshed leaderboard at 84.5, ahead of GPT-5.5’s 80.2; Google’s own Gemini models trail both.
- 01Google updates Android Bench with new LLMs, but Gemini still lags behindarstechnica.com · reporting
