brief researchproducts
Google overhauls its coding-agent benchmark
Android Bench switched to the Harbor evaluation framework and added eight models; Claude Fable 5 currently leads at 84.5 versus GPT-5.5's 80.2.
Google overhauled its Android Bench coding-agent benchmark on July 8, switching its evaluation harness to the standardized “Harbor” framework and adding eight new models, including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, and two Qwen variants. Claude Fable 5 currently leads the refreshed leaderboard at 84.5, ahead of GPT-5.5’s 80.2; Google’s own Gemini models trail both.
sources 1 cited
1 arstechnica.com Google updates Android Bench with new LLMs, but Gemini still lags behind