Google releases Android Bench 2.0 with long-horizon AI coding tasks
New benchmark evaluates LLMs and agents on multi-day engineering challenges with continuous scoring instead of binary pass/fail.
What to know
- Google's Android Bench 2.0 introduces long-horizon tasks simulating multi-day engineering work, with pass rates around 28% versus 91% on simpler original tasks.
- New continuous scoring system replaces binary pass/fail grading to better reflect partial progress on complex challenges.
- Analysis shows AI excels at writing new code and deterministic transformations (Java to Kotlin) but struggles with refactoring, migrations, and cross-platform app porting.
Google Android team Benchmark creator and publisher
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Developer notes GPT Sol and Astra lead benchmark, awaits Swift evaluation
Kelset, a developer on Mastodon, comments that the benchmark results align with their own experience, noting GPT Sol and Astra (along with Fable) rank at the top. They express interest in similar evaluation benchmarks for Swift and SwiftUI development.
“not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)…”
— kelset@mastodon.online -
K
The Android team released v2 of their benchmark for LLMs: https:// android-developers.googleblog. com/2026/09/android-bench-2-long-horizon-tasks.html not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)
-
-
1
Android Bench 2.0 reveals significant performance gap on long-horizon tasks
Results show the highest pass rate for long-horizon tasks is around 28%, dramatically lower than the ~91% pass rate on the original benchmark's simpler tasks. Analysis reveals AI excels at writing new code but struggles with refactoring and migrations due to architectural complexity, and cross-platform app porting remains unsolved by all tested models.
“The highest pass rate for LHTs is around 28%, much lower than the ~91% for the original tasks in the benchmark.”
— Google Android team -
2 outlets first by blog.google, 9d ago · also 9to5Google · read ↗
-
-
background
Google shifts from binary to continuous scoring for complex tasks — Google moves away from binary pass/fail grading toward continuous scoring that better captures partial progress on multi-day tasks. Example: an agent that refactors 40 screens to Jetpack Compose, sets up database tables, and passes 90% of requirements but fails one edge-case assertion would score 0% under binary grading but receives a meaningful completion rate under the new system.
-
background
Google releases Android Bench 2.0 with long-horizon task evaluation — Google announces Android Bench 2.0, introducing long-horizon tasks (LHTs)—complex multi-day engineering challenges like building apps from scratch, upgrading dependencies, adding features, and converting cross-platform apps to Android. The update also introduces agentic evaluation for testing AI agents from model providers.
Also covered reported alongside — the timeline has no entry for these yet
-
first by NewsBytes, 9d ago · also Android Headlines
1 more headline
- Google’s Android Bench 2.0 Replaces Pass/Fail Grades for Real-World Coding Tests Android Headlines · 9d ago