
Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Quick Answer
Google's Android Bench 2.0 introduces long-horizon tasks and agentic evaluation for AI model assessment, enhancing scoring nuances.
Quick Take
The update reveals that AI excels in writing new code but struggles with complex refactoring, with top models like Claude Opus 5.5 achieving a 32% LHT pass rate.
Key Points
- Android Bench 2.0 adds support for long-horizon tasks (LHTs) and agentic evaluation.
- The new scoring system offers nuanced assessments beyond binary pass/fail.
- AI performs better at writing new code than refactoring existing code.
- Claude Opus 5.5 leads with a 32% LHT pass rate, followed by GPT 6 Astra at 28%.
- Challenges remain for AI in tasks requiring runtime validation and complex migrations.
DeepSignal Analysis
What happened
Google has launched Android Bench 2.0, enhancing its benchmark framework for AI model evaluation in Android development. This update introduces long-horizon tasks (LHTs) and a continuous scoring system, moving away from binary evaluations. The top-performing model, Claude Opus 5.5, achieved a 32% pass rate on LHTs, indicating varying levels of AI capability in coding tasks.
Key evidence
- Android Bench 2.0 introduces long-horizon tasks, which are complex tasks that can take engineers multiple days or even a week to complete.
- The updated scoring system allows for a nuanced evaluation, where a task can be partially successful rather than simply pass or fail.
- Claude Opus 5.5 leads the benchmark with a 32% pass rate on long-horizon tasks, while GPT 6 Astra follows with a 28% pass rate.
Why it matters
The introduction of long-horizon tasks and a continuous scoring system reflects a significant advancement in how AI models are evaluated in software development. This shift allows for a more accurate assessment of AI capabilities, particularly in complex coding scenarios. Understanding these capabilities can guide developers in choosing the right AI tools for specific tasks, potentially improving productivity and code quality.
📖 Reader Mode
~3 min readGoogle has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess performance on complex, multi-step development tasks.
Launched a few months ago, Android Bench evaluates AI models against a set of common development tasks, incorporating Android best practices in areas such as permissions, navigation, and connectivity. With the latest release, Google has expanded the benchmark to cover a broader range of development tasks.
Today we're releasing the first set of long-horizon tasks (LHT), which are tasks of great complexity that take an engineer multiple days or even a week to complete. We are also introducing agentic evaluation, starting with agents from corresponding model providers.
While the original version focused on incremental changes to existing repositories, Android Bench 2.0 introduces long-horizon tasks (LHTs) that Google describes as work an engineer might take "multiple days or even a week to complete". These tasks include upgrading dependencies, adding new features, building apps from scratch, or converting a cross-platform app to Android.
One major change in version 2.0 is the shift from binary pass/fail evaluation to a more nuanced scoring system. Previously, a complex task could be marked as failed because of a single failing edge-case assertion, even if the agent had successfully met dozens of other requirements.
We calculate this completion rate through a combination of factors like functionality, visual fidelity, and avoiding regressions. We also apply objective scoring penalties for deviations from evaluation instructions or structural constraints.
The results provided by Android Bench 2.0 also help identify which tasks are more likely to succeed with AI assistance. For example, Google reports that "AI does a better job at writing new code rather than refactoring existing code", which can reflect the fact that refactoring and migrations require an understanding of the architectural complexity of the codebase.
Similarly, AI performs well on several "well-established, deterministic transformations", even in larger codebases. Examples include converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer.
On the contrary, in a number of cases model still struggle, including with tasks requiring runtime validation (like missing dependency injection graphs), involving breaking framework changes, or running into knowledge gaps with unreleased libraries. In particular, the best-in-class model achieves only 80% completion rate in porting cross-platform app to Android, which remains overall an "open challenge".
The updated Android Bench 2.0 dashboard includes recent models, such as Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. At the time of the article, Claude Opus 5.5 was at the top of the leaderboard with a 32% LHT pass rate, followed by GPT 6 Astra at 28%.
About the Author
Sergio De Simone
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

