Google has updated Android Bench, a benchmark for evaluating large language models (LLMs) in Android app development, adding eight new models including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. The update also introduces metrics for cost and efficiency, alongside open-weight models, and adopts a new framework designed to be easier for developers to use. Developers are invited to run their own tests and submit feedback to help shape the benchmark's future.
Google revamps Android Bench with new LLMs and improved evaluation framework
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments