Google has updated its Android Bench benchmark, which evaluates large language models on Android development tasks. The update adds eight new models, including Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max.
The benchmark now incorporates cost and efficiency metrics, along with a more user-friendly framework. Developers are encouraged to run their own tests and provide feedback to help shape future versions.
Comments