Google has updated its Android Bench benchmark, which evaluates large language models on 100 Android development tasks, by adding eight new models including Claude Fable 5, Claude Sonnet 5, and Qwen 3.7 Max. The update also introduces a new framework designed for easier use and incorporates metrics for cost and efficiency alongside open-weight models. Developers are encouraged to run their own tests and submit feedback to help shape the benchmark's future.
Google Refreshes Android Bench With New LLMs and a Simpler Testing Framework
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments