Google has updated its Android Bench benchmark, adding eight new large language models including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. The benchmark measures how well AI agents perform on 100 Android development tasks and now includes metrics for cost and efficiency. Developers can run their own tests and submit feedback to help shape future versions of Android Bench.
Google revamps Android Bench benchmark with eight new language models
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments