← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Google releases Android Bench 2.0 for evaluating AI on complex development tasks

🔄 Updated 6d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Android Bench 2.0 evaluates AI on complex tasks like building apps and adding features.
  • The new version uses continuous scoring instead of binary pass/fail.
  • GPT-6 Astra achieved the highest pass rate at 28% on the new benchmark.
  • AI models struggle with refactoring and tasks requiring runtime validation.

Android Bench Evolves for Complex Tasks

Google has updated its Android Bench to version 2.0, shifting its focus to evaluate AI models on more complex and time-consuming development tasks. The initial version concentrated on incremental changes and bug fixes within existing codebases. Android Bench 2.0 now targets tasks that could take an engineer multiple days or weeks, such as developing new features, creating applications from scratch, and converting cross-platform applications to Android.

Continuous Scoring for Nuanced Evaluation

To accommodate the increased complexity of these tasks, Android Bench 2.0 introduces a continuous scoring system, moving away from simple pass or fail grades. This new system calculates a completion rate based on factors like functionality, visual fidelity, and the absence of regressions. Objective scoring penalties are applied for deviations from instructions or structural constraints, allowing for a more detailed assessment of AI performance.

Model Performance and Challenges

Several AI models, including Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max, have been rated using the new benchmark. GPT-6 Astra currently leads with a 28% pass rate. Google's findings indicate that porting cross-platform apps to Android remains a significant challenge for AI, with no model achieving a 100% pass rate and frontier models reaching a maximum of 80% completion. AI models perform better at writing new code than refactoring existing code, and struggle with tasks requiring runtime validation or involving breaking framework changes.

Agent Evaluations and Future Plans

Android Bench 2.0 also incorporates agent evaluations, running agents from corresponding model providers, such as Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex. Google notes that "harness design positively impacts developer outcomes." The benchmark plans to include various model and agent combinations in future evaluations to further explore their capabilities in software development.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~26 min · 21 stories · Sep 23

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Google released Android Bench 2.0, an updated benchmark designed to evaluate AI models on more complex, long-horizon Android development tasks, moving beyond simple bug fixes to include feature additions and app conversions. This update introduces continuous scoring to assess functionality, visual fidelity, and regressions, providing a more nuanced evaluation of AI agent capabilities in software development.