Logic of Logic
thursday, august 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
brief researchproducts

Google overhauls its coding-agent benchmark

Android Bench switched to the Harbor evaluation framework and added eight models; Claude Fable 5 currently leads at 84.5 versus GPT-5.5's 80.2.

Google overhauled its Android Bench coding-agent benchmark on July 8, switching its evaluation harness to the standardized “Harbor” framework and adding eight new models, including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, and two Qwen variants. Claude Fable 5 currently leads the refreshed leaderboard at 84.5, ahead of GPT-5.5’s 80.2; Google’s own Gemini models trail both.

sources 1 cited
1 arstechnica.com Google updates Android Bench with new LLMs, but Gemini still lags behind
next