AI
Google releases Android Bench 2.0 with long-horizon AI coding tasks
Android Bench 2.0 reveals significant performance gap on long-horizon tasks
1 Sep 17 12:00 PM · 9d ago · 4 articles · 1 source · development 1 of 2
Results show the highest pass rate for long-horizon tasks is around 28%, dramatically lower than the ~91% pass rate on the original benchmark's simpler tasks. Analysis reveals AI excels at writing new code but struggles with refactoring and migrations due to architectural complexity, and cross-platform app porting remains unsolved by all tested models.
“The highest pass rate for LHTs is around 28%, much lower than the ~91% for the original tasks in the benchmark.”
Google Android teamGoogle Android team Benchmark creator and publisher
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 4:30 PM ET
What was reported 1 claim about this development
-
first by blog.google, 9d ago · also 9to5Google
1 more headline
- Android Bench 2.0 focuses on long-horizon tasks, agent evaluations 9to5Google · 9d ago
All 2 developments of Google releases Android Bench 2.0 with long-horizon AI… →
Google NewsMastodonHacker News