Google has updated its Android Bench benchmark, which evaluates large language models (LLMs) on 100 Android development tasks, adding eight new models including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. The update also introduces metrics for cost and efficiency, as well as open-weight models. Developers can now run their own tests and submit feedback to help shape the benchmark's future.
Google refreshes Android Bench benchmark with eight new AI models
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments