Android Bench 2.0 adds harder Android tasks and continuous scoring
Google has changed how Android AI work is measured, replacing simple pass/fail checks with scoring for more complex development tasks.
Google has updated Android Bench to version 2.0, adding long-horizon tasks, agent-based evaluation, and continuous scoring so AI systems can be judged on harder Android development work. The new benchmark now covers jobs such as dependency upgrades, new feature work, building apps from scratch, and converting cross-platform apps to Android, which are the kinds of tasks that can take days rather than minutes. It also replaces simple pass/fail grading with scoring that weighs functionality, visual quality, regressions, and instruction-following penalties. In the dashboard Google highlighted that Claude Opus 5.5 led the long-horizon tasks leaderboard at 32% pass rate, with GPT 6 Astra at 28%, and said the best model only reached 80% on cross-platform app porting. The release is meant to show where AI is useful in development now, especially for new code and deterministic migrations, and where it still falls short on refactors, runtime validation, and framework changes.
Why it matters
The update makes the benchmark better suited to work that takes longer than a quick code change, including app builds, dependency upgrades and app porting. That gives teams a clearer picture of where models can help on Android development today and where they still leave work unfinished, especially on harder migration tasks.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- InfoQ