Google Overhauls Android Bench Benchmark with Eight New LLMs and Improved Framework

Google has updated its Android Bench benchmark, originally launched in March, to evaluate large language models on Android development tasks. The update adds eight new models, including Claude Fable 5, Claude Sonnet 5, and Qwen 3.7 Max, along with new metrics for cost and efficiency.

A new, more user-friendly framework allows developers to run their own tests and submit feedback to shape the benchmark's future. The leaderboard now also includes open-weight models, aiming to help developers choose the best AI agent for coding tasks.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleWeb Scraper Declares 'Google and Reddit Do Not Own the Internet' After Court Victory
Start typing to search