Google has updated its Android Bench to version 2.0, shifting its focus to evaluate AI models on more complex and time-consuming development tasks. The initial version concentrated on incremental changes and bug fixes within existing codebases. Android Bench 2.0 now targets tasks that could take an engineer multiple days or weeks, such as developing new features, creating applications from scratch, and converting cross-platform applications to Android.
To accommodate the increased complexity of these tasks, Android Bench 2.0 introduces a continuous scoring system, moving away from simple pass or fail grades. This new system calculates a completion rate based on factors like functionality, visual fidelity, and the absence of regressions. Objective scoring penalties are applied for deviations from instructions or structural constraints, allowing for a more detailed assessment of AI performance.
Several AI models, including Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max, have been rated using the new benchmark. GPT-6 Astra currently leads with a 28% pass rate. Google's findings indicate that porting cross-platform apps to Android remains a significant challenge for AI, with no model achieving a 100% pass rate and frontier models reaching a maximum of 80% completion. AI models perform better at writing new code than refactoring existing code, and struggle with tasks requiring runtime validation or involving breaking framework changes.
Android Bench 2.0 also incorporates agent evaluations, running agents from corresponding model providers, such as Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex. Google notes that "harness design positively impacts developer outcomes." The benchmark plans to include various model and agent combinations in future evaluations to further explore their capabilities in software development.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Google released Android Bench 2.0, an updated benchmark designed to evaluate AI models on more complex, long-horizon Android development tasks, moving beyond simple bug fixes to include feature additions and app conversions. This update introduces continuous scoring to assess functionality, visual fidelity, and regressions, providing a more nuanced evaluation of AI agent capabilities in software development.